Let AI Interpret Color, Let NLP Render

Hi Nate and everyone,

I recently did a small experiment with AI image editing that gave me a result I honestly wasn’t expecting, and it made me wonder whether there might be an interesting application for this kind of technology in Negative Lab Pro.

This is not a suggestion to add generative AI image editing to NLP. In fact, the experiment convinced me that generative image output is exactly what I wouldn’t want in a negative conversion workflow.

The experiment

I camera-scan my negatives as RAW files. For this test, I took a C-41 negative and did essentially the minimum possible processing:

Negative → simple RGB inversion → linear tone curve

I deliberately did not correct the orange mask, white balance, RGB channels, contrast, black/white points, or color.

The resulting positive was therefore extremely cyan/blue and looked nowhere near a finished scan.

I then gave this image to a modern vision/generative AI model and simply asked it to make the exposure and colors look natural and accurate.

To my surprise, the AI was able to infer a very convincing color balance and tonal interpretation from this extremely poor starting point. It seemed capable of using contextual information in the scene, such as pavement, vegetation, skin tones, neutral objects, vehicles, daylight, etc., to infer what a plausible color rendering should be.

I also tried another experiment with a photograph of some objects whose real-world colors I know very well. I compared an NLP conversion with the AI interpretation and showed both versions to someone familiar with the actual scene without telling them which was which. They actually preferred the AI version as being closer to how the objects looked to the eye.

That was the part that really surprised me.

But there is a major problem

The AI result looks excellent at normal viewing size, but when inspected at 100%, the weakness becomes obvious.

Because the model is generating/reconstructing the image rather than processing the original pixels deterministically, it can subtly change fine details. Text and license plates may become incorrect, object edges may become slightly too clean, grain structure can change, and very fine details can be reconstructed rather than preserved.

For photography, especially archival work, printing, or professional output, I obviously don’t want that.

And that led me to a different thought:

What if we separate AI “understanding” from AI “generation”?

Instead of asking an AI model to output an image, could a vision model be used only as a scene/color interpretation engine?

For example, it might analyze an inverted negative and estimate things such as:

  • likely illuminant / white balance
  • reliable neutral references
  • overall color cast
  • plausible black and white points
  • RGB channel relationships
  • exposure and contrast targets
  • skin-tone references where appropriate
  • confidence values for each interpretation

The AI would never generate or modify pixels.

It would only provide parameters or targets to NLP. NLP’s existing deterministic negative-processing/rendering pipeline would then perform the actual conversion using the original image data.

In other words:

Let AI interpret. Let NLP render.

This seems interesting to me because the two technologies appear to have almost complementary strengths.

NLP is excellent at preserving the photographic data and producing a deterministic conversion from the original negative. Modern vision models, on the other hand, seem remarkably good at understanding the semantic color relationships of a scene, even when presented with an image that has a very severe color cast.

Combining those two abilities might potentially provide some of the scene intelligence of modern AI without introducing any generative artifacts.

An even more interesting possibility: roll-level analysis

Since NLP already works naturally with rolls of film, perhaps such a system wouldn’t even need to analyze every frame independently.

A vision model could analyze multiple frames from the same roll, identify the most reliable neutral references, daylight scenes, skin tones, etc., establish a general interpretation for that roll, and then allow NLP to apply frame-specific corrections on top of it.

That could potentially preserve consistency across a whole roll while still benefiting from scene-aware color interpretation.

Of course, I’m describing this from a photographer/user perspective rather than claiming to know how practical it would be within NLP’s architecture. There would obviously be many questions around model size, performance, training data, reliability, local vs. cloud inference, and preventing semantic assumptions from overriding actual film data.

In particular, I think the original negative data should always have the final say. If an AI thinks an object “should” be neutral gray but the negative clearly says it is green, the system should trust the negative, not the AI’s world knowledge.

But after seeing how well a vision model could infer natural color from an intentionally terrible linear inversion, I thought the idea was interesting enough to share.

I’d be very curious to hear what Nate thinks about this. Is something along these lines technically plausible for a future NLP workflow, or are there fundamental reasons why it wouldn’t make sense compared with the current approach?

And don’t worry, I’m not asking you to delay Standalone for another year. :grinning_face_with_smiling_eyes:

1 Like

I suppose that we’ll see AI conversions done in or instead of existing apps rather sooner than later. Whether these conversions will be a) true or b) fit our expectations remains to be seen. Interesting times ahead!

1 Like

Interesting.

Could you provide some more info about this?

  1. Can you specify the model you used, as well as any model parameters you controlled e.g. effort level and such?
  2. What was the model input - JPEG, TIFF, something else?
  3. How many images have you processed this way?
  4. How was the consistency across multiple images?
    1. Did you try different films?
      1. Was a roll of film consistent?
    2. How about different lighting conditions?

These are all good points, but one major obstacle remains. I agree it would be nice to have AI infer the colors, not even all colors but major colors and have all other colors fall on line. When one looks at a picture , she/he subconsciously evaluate, pavement, greenery, sky , common logos like McDonald or street signs and makes the judgement how well colors turned out. So yes, theoretically software can run analysis and infer what correction needs to be applied to get those major colors properly rendered. If the curves are “smooth” enough the remaining colors like skin tones will fall into reasonable range. The obvious disadvantage of that approach that all pictures will look alike with all subtle shifts in colors due to actual light conditions, specific film stock and alike will be eliminated. The beauty of film, specifically slide film, is that it captures the scene the way even our eyes cannot see it - in scientifically correct way with regard of color temperature and reflective properties of the elements. Our task mostly is to recover that capture - while possibly correcting overall tint and color shifts of the resulting image to match the viewing conditions under which image is seen. As for similar attempt to use AI to invert the negative see my article https://medium.com/full-frame/i-asked-ai-to-invert-film-negative-99073a565d1f?sk=c149b7aa006b15bbadab0a3874cad80d which generated much more views than I ever anticipated

1 Like

The reversed text was automagically flipped to be the correct way in two instances and in the fake image both these ‘corrected’ texts were the same colour whereas in the original they appear to be different. Regardless of how ‘good’ the resultant image may look it was still wrong. What is the point of taking a photo of something then risking its destruction by letting a turbo charged search engine loose on it. Ai has its uses, but people are getting lazy. Half my enjoyment of taking photos it in the post processing, I’ll do a straight edit then for my own enjoyment I’ll make alternative edits to give the image a different look. Sure these alternatives aren’t depicting the image as it was when shot but makes for an interesting look. Still each to their own as long as they admit it was AI rather than their own creative input.

I think we may be talking about two slightly different things here. I completely agree that an AI-generated image is not a faithful representation of the original photograph. It is generative by nature, and I would not regard that generated result as the photograph itself.In fact, the altered text, reconstructed detail and changed pixels are exactly why I said in my original post that generative image output is not what I would want in a negative conversion workflow.What I am interested in is a different part of the process: whether AI can correctly identify the semantic and colour relationships within an image, and whether that information can then be used to give us a sensible starting point for colour correction and curve adjustments.The attached image shows what I mean. The system is identifying separate elements in the scene, such as purple pansies, green foliage, a red bin, a black lamp post, the stone planter and the metal railing. That kind of scene and colour recognition is the part I am exploring.From there, the actual correction would still be performed on the original image data. The AI would only provide guidance, targets or parameters for the curve, while NLP would still do the actual rendering. The final colour, contrast and stylistic decisions would remain under the photographer’s control. In other words, I am not suggesting that generative AI should replace or reconstruct the original pixels.And importantly, the negative itself should always have the final say. If the AI’s semantic assumption conflicts with what is actually present in the image data, then the image data should win.Tools such as Negative Lab Pro already give us an algorithmic starting point rather than requiring us to invert the negative and construct every curve entirely by hand. Hasselblad’s Phocus Mobile uses a similar principle with HNNR, where sampled information is used to constrain and guide the curve rather than generate replacement content.So I do understand your concern about generative AI, and I agree with it in that context. I think we actually agree on that point. The distinction I am trying to make is between using AI to analyse and understand the image and using AI to generate a new one. Those are quite different applications.:blush:

I’d say don’t non-deterministically solve what can be solved deterministically. The perceived randomness and effort in post-processing is largely attributed to not scanning color negative stock optimally in the first place. Why reaching out for AI to “interpret” color if we didn’t even bother in scanning to the best of our abilities?

We have plenty of scientific research. We have decades old Frontier machines that tell a story. We have the most expensive scanners bringing us The Odyssey to our screens as we speak. They all have something in common. They don’t scan with white backlight hoping for AI to rescue color appearance.

Yes, the wheel doesn’t need to be invented. It has been invented decades ago. It’s just that the ingredients only now become available to us as regular consumers (sequential narrowband capable hardware and software that is).

Color negative derived work will always remain interpretive at its core. But for sure let us start from higher ground and not throw the sledgehammer at it because we run out of options. We are not.

Of course you are correct - the wheel has been invented long ago, it’s just the wheel is very expensive and slow turning :wink:

What people want is simplicity, cost and speed of Kodak Scanza and quality of Noritsu or Lasergraphics. and we already know that’s sort of possible if any of these requirements are dropped.

Let’s just recognize that we are trying to use device (digital camera) which has been optimized for completely different task.

Everybody in the field know that simplest working solution is monochrome camera, RGB (three or four if IR to be included) capture and appropriate software. How many people are willing to go this route - I would speculate only 3-5% of folks are ready for this . Not only this system will cost between $1000 and $2000 dollars - as now one have to acquire the camera - but the learning curve is still steep.

Everybody else will continue use of the shelf digital camera, cs-lite and NLP and they will be content with the quality already available which is not bad if people don’t go crazy and use indiscrimantly push and pull, try to make a candy from shitty film expired 50 years ago and use completely exhausted developer to save a buck.

We know that properly exposed and developed film contribute to 75% of success with remaining 25% coming from knowing scanning routine well enough.

We know not to expect anything good from scans which have perforation holes and edge marking included - but having these “true film” attributes is a matter of pride for huge number of people.

Sorry for my long rant, my point is that yes we can already get decent results with tools we already use. The matter is that field is under pressure from newbies driven by pure vanity who want smart phone quality pictures from the media which was designed for much more thoughtful and patient customers in mind.

3 Likes

Scanning with a more optimal approach doesn’t need to be expensive.

  • Suitable backlight: 300-400$
  • Suitable Software: Free
  • Suitable Camera: Existing DSLR

Learning Curve is modest as well. If calibrating once per film and then drag-n-dropping 3 files instead of 1 is rewarded with more predictable output then the trade-off might be justified in comparison for quite some of us.

As this thread is about using AI for interpreting color I would just steer away from making a problem more complicated first and then solve it with AI.

I get that you are throwing the idea of AI just examining the image to try and get the colours, luminance Saturation etc but my response was based on the image on the linked page.
No AI will know what the ambience of the scene was at the moment of capture. What was the light like - was it cool, warm, affected by some nearby light source causing a colour cast.

The AI may be able to determine which flowers are in a scene or be able to recognise a store logo like McDonalds / burger king & colour them as the correct colour but what if the scene due to surrounding influence looked different? The correct colour under ideal light may not be the colour of the scene.

I don’t rely on photography for a living - if I did I would probably be a pauper. I would rather do the whole post processing myself with some help from decent editing software. AI still isn’t anywhere near what they would have us believe. Yes I can agree that it has come on leaps & bounds but a lot of the stuff we see and hear that ai has ‘created’ is blatantly obvious that it was made by ai and is vastly inferior to what a real living breathing human could have produced. I do use ai but only on subjects that have many sources, AI can find these sources faster than anyone could then summarise the results. Most AI’s also show that disclaimer that the answers may be incorrect. This disclaimer says a whole lot more than I ever could.

AI is here to stay but it should always remain an opt in rather than opt out or worse, forced upon us. I’m fine with it being something you cherish - it’s just not for everyone.