Let AI Interpret Color, Let NLP Render

Hi Nate and everyone,

I recently did a small experiment with AI image editing that gave me a result I honestly wasn’t expecting, and it made me wonder whether there might be an interesting application for this kind of technology in Negative Lab Pro.

This is not a suggestion to add generative AI image editing to NLP. In fact, the experiment convinced me that generative image output is exactly what I wouldn’t want in a negative conversion workflow.

The experiment

I camera-scan my negatives as RAW files. For this test, I took a C-41 negative and did essentially the minimum possible processing:

Negative → simple RGB inversion → linear tone curve

I deliberately did not correct the orange mask, white balance, RGB channels, contrast, black/white points, or color.

The resulting positive was therefore extremely cyan/blue and looked nowhere near a finished scan.

I then gave this image to a modern vision/generative AI model and simply asked it to make the exposure and colors look natural and accurate.

To my surprise, the AI was able to infer a very convincing color balance and tonal interpretation from this extremely poor starting point. It seemed capable of using contextual information in the scene, such as pavement, vegetation, skin tones, neutral objects, vehicles, daylight, etc., to infer what a plausible color rendering should be.

I also tried another experiment with a photograph of some objects whose real-world colors I know very well. I compared an NLP conversion with the AI interpretation and showed both versions to someone familiar with the actual scene without telling them which was which. They actually preferred the AI version as being closer to how the objects looked to the eye.

That was the part that really surprised me.

But there is a major problem

The AI result looks excellent at normal viewing size, but when inspected at 100%, the weakness becomes obvious.

Because the model is generating/reconstructing the image rather than processing the original pixels deterministically, it can subtly change fine details. Text and license plates may become incorrect, object edges may become slightly too clean, grain structure can change, and very fine details can be reconstructed rather than preserved.

For photography, especially archival work, printing, or professional output, I obviously don’t want that.

And that led me to a different thought:

What if we separate AI “understanding” from AI “generation”?

Instead of asking an AI model to output an image, could a vision model be used only as a scene/color interpretation engine?

For example, it might analyze an inverted negative and estimate things such as:

  • likely illuminant / white balance
  • reliable neutral references
  • overall color cast
  • plausible black and white points
  • RGB channel relationships
  • exposure and contrast targets
  • skin-tone references where appropriate
  • confidence values for each interpretation

The AI would never generate or modify pixels.

It would only provide parameters or targets to NLP. NLP’s existing deterministic negative-processing/rendering pipeline would then perform the actual conversion using the original image data.

In other words:

Let AI interpret. Let NLP render.

This seems interesting to me because the two technologies appear to have almost complementary strengths.

NLP is excellent at preserving the photographic data and producing a deterministic conversion from the original negative. Modern vision models, on the other hand, seem remarkably good at understanding the semantic color relationships of a scene, even when presented with an image that has a very severe color cast.

Combining those two abilities might potentially provide some of the scene intelligence of modern AI without introducing any generative artifacts.

An even more interesting possibility: roll-level analysis

Since NLP already works naturally with rolls of film, perhaps such a system wouldn’t even need to analyze every frame independently.

A vision model could analyze multiple frames from the same roll, identify the most reliable neutral references, daylight scenes, skin tones, etc., establish a general interpretation for that roll, and then allow NLP to apply frame-specific corrections on top of it.

That could potentially preserve consistency across a whole roll while still benefiting from scene-aware color interpretation.

Of course, I’m describing this from a photographer/user perspective rather than claiming to know how practical it would be within NLP’s architecture. There would obviously be many questions around model size, performance, training data, reliability, local vs. cloud inference, and preventing semantic assumptions from overriding actual film data.

In particular, I think the original negative data should always have the final say. If an AI thinks an object “should” be neutral gray but the negative clearly says it is green, the system should trust the negative, not the AI’s world knowledge.

But after seeing how well a vision model could infer natural color from an intentionally terrible linear inversion, I thought the idea was interesting enough to share.

I’d be very curious to hear what Nate thinks about this. Is something along these lines technically plausible for a future NLP workflow, or are there fundamental reasons why it wouldn’t make sense compared with the current approach?

And don’t worry, I’m not asking you to delay Standalone for another year. :grinning_face_with_smiling_eyes:

I suppose that we’ll see AI conversions done in or instead of existing apps rather sooner than later. Whether these conversions will be a) true or b) fit our expectations remains to be seen. Interesting times ahead!

1 Like

Interesting.

Could you provide some more info about this?

  1. Can you specify the model you used, as well as any model parameters you controlled e.g. effort level and such?
  2. What was the model input - JPEG, TIFF, something else?
  3. How many images have you processed this way?
  4. How was the consistency across multiple images?
    1. Did you try different films?
      1. Was a roll of film consistent?
    2. How about different lighting conditions?