The Latent Space Technique Adobe (and Everyone Else) Will Steal Next (Infinite Crop/Zoom using semantic context awareness).
No, this is not another “infinite zoom” workflow.
We’ve had infinite zooms, recursive img2img, outpainting and endless upscaling for years. They all eventually hit the same wall: the deeper you zoom, the less the model understands what it is looking at. It stops seeing an eagle’s eye and starts seeing a brown circle with random textures. Consistency slowly disappears and the generated details become generic noise.
The problem isn’t image quality.
The problem is semantics.
If you’ve read my previous articles, you’ll probably notice a recurring theme. Whether I was talking about infinite scene generation, consistent comics, prompt randomization or Krea 2 workflows, I kept coming back to the same conclusion:
Semantic understanding is far more important than pixel similarity.
Read MoreFrom Infinite Scene Images to Infinite Comic Books (ComfyUI first comic book generator from a simple story with consistancy using no references, LORAs etc).
A few days ago I published an article about a workflow capable of generating an effectively infinite number of consistent scene images using nothing more than a text description. The central idea was surprisingly simple: instead of maintaining consistency by carrying visual information from one generation to the next through LoRAs, ControlNet, reference images, edit models or image-to-image workflows, I continuously regenerated the description of the world itself. The previous image stopped being the source of truth. The description became the source of truth.






If you haven’t read it yet, you can find it here:
KREA 2 (and maybe others) infinite scene images with consistancy using description (Comfy UI workflow)
Krea 2 (and probably other image models): Infinite Character Consistency Using Descriptions + ComfyUI
Ok, so long story short, I saw a few workflows that enabled Krea 2 inside ComfyUI to generate a single image with four panels in order to keep character consistency. Nice idea, but for me there was one problem: resolution. Four panels inside one image is a no-go for the kind of work I do.
Then I had a different idea.
What if Krea is simply good enough that, with sufficiently detailed descriptions, it doesn’t need previous images at all? What if every scene completely recreates the characters from scratch?
So I tested it.
It works surprisingly well.
The characters remain remarkably consistent across an essentially unlimited number of separately generated images.
The workflow is at the end of the article.
Read MoreAce Step 1.5 XL ComfyUI workflow for generating random tags, generate song and then give it a rating by using waveform analysis
The idea came to me after sorting trough a lot of Ace Step 1.5 XL outputs and trying to find best styles and tags for songs. Why not automate the generation process AND the review process, or at least make it easier. So as usual I used Qwen LM and Qwen VL (compared to something like olama these ones run directly in comfy and do not require a server) to randomize the tags on each run, but more importantly to try and rate the output. How ? By converting the audio output into a set of waveforms for 4 segments of the song that I feed into Qwen VL as an image and ask it to subjectively look at the waveform and give it feedback and rating, rating that is used then to also name the output file. Like this. I am not sure it works properly but the A+ rated songs were indeed better than B rated ones.
Workflow is here. Install the missing extensions and add the qwen models.



