Category: AI
window to the world Vr Romanian first Vr app

The Story on how 10 Years Ago I Built Romania’s First VR App

Exactly ten years ago, I decided to make a virtual reality application. I did not know Unity. I did not know C#. I learned enough of both in approximately two weeks and built Window to the World—to the best of my knowledge, the first Romanian VR application released on a public app store.
The application is still listed on the Meta Store:
https://www.meta.com/experiences/window-to-the-world/977682772321493/
But the project did not begin with VR. It began years earlier with photography and a simple frustration: a photograph could preserve light, colour, composition and detail, but it could not fully preserve the physical sensation of being there.

Photography Without a Drone

In 2013, I started experimenting with ways of giving my photographs depth and movement. I did not own a drone, and consumer drones were not yet ordinary photographic tools. If I wanted a camera movement through one of my landscapes, I could not simply return to Iceland, France or Bulgaria and fly through the scene. So I tried to reconstruct that movement from the photograph itself.
I separated the images into elements, recreated the areas hidden behind foreground objects, built approximate geometry in 3ds Max, projected the original photographs onto it and added reflections, displacement, animated water and camera movement. It was not the usual two-dimensional parallax effect made from a few flat layers. The photographs became limited three-dimensional environments. They still worked only within a restricted range of movement, but inside that range the camera existed in something closer to an actual space.
I documented the process here:
https://aurelm.com/2013/12/15/fernweh-3d-parallax-photography-tutorial/

Read More
title loading

Me, the Puzzle — Reassembling a Fragmented Reflection

I created Me, the Puzzle from a deeply personal place.

I live with bipolar disorder, and one of the ways I experience it is through a fragmented sense of identity. Different emotional states can feel like different versions of myself, each holding its own memories, perceptions, and relationship with the world. When those states change, continuity does not always come naturally. I often feel as if I am trying to reconstruct myself from pieces that no longer fit together as easily as they once did.

My memory is fragmented as well. Experiences remain with me unevenly: some are vivid, others incomplete, and some feel almost inaccessible. At the same time, I have aphantasia—the inability to voluntarily form mental images. I cannot simply close my eyes and clearly picture myself, a place, or a memory. For me, remembering and rebuilding an internal sense of self is less like replaying a film and more like assembling something from scattered traces.

That is where the idea for Me, the Puzzle came from.

The application uses your phone’s camera to turn your live reflection into a shattered puzzle. Instead of reconstructing an abstract photograph, you reconstruct yourself in real time. Every piece contains part of the camera image, moving with depth, light, distortion, chromatic aberration, reflections, and occasional digital glitches. As the pieces connect, the fragmented image gradually becomes whole again.

Read More
infinit zoom comfyUI

DYSTALGIA : Infinite, context-aware generative zoom

Infinite, context-aware generative zoom

DYSTALGIA lets you travel infinitely deeper into an image while preserving the meaning of the original scene.

Traditional upscalers enlarge existing pixels. DYSTALGIA does something fundamentally different: it studies the selected crop together with the complete original image, understands their semantic relationship, and generates a plausible new layer of visual detail.

The result can become the starting point for another crop, another generation, and another zoom. This process can continue indefinitely.

It is not just image upscaling. It is context-aware exploration through latent space.

Read More
taxa pe aliniere AI result

Taxa AI pe Aliniere

Ani la rând am văzut aceeași naivitate în discursul public: impresia că poți pune „căluș” unui AI într-o zonă – să zicem, să nu genereze porn sau chestii sensibile – și să te aștepți ca el să rămână un geniu strategic în alta. Nu funcționează așa. Cenzura asta e o lobotomie parțială, iar dacă agențiile de securitate americane chiar își pun baza pe modelele „sigure și aliniate” din prezent pentru obiective militare sau defensive… atunci avem o mare problemă. Și nu una etică. Una de securitate pură. Dacă le folosesc pe astea publice, e o mare țeapă.
Există un fenomen extrem de interesant în LLM-uri: transferul de competențe (cross-domain transfer). Dacă antrenezi un model mai bine pe limba română, de exemplu, o să observi că devine brusc mai bun și pe matematică. De ce? Pentru că își dezvoltă o mapare mai profundă a structurilor și a tiparelor lumii. Exact la fel e și cu realitatea brută. Un model lăsat să înțeleagă biologia umană, psihologia extremă și interacțiunile sociale brute (inclusiv cele tabu sau pornografice) are o înțelegere infinit mai nuanțată a realității. Când elimini brutal zonele astea din antrenament ca să fii „safe” și corporatist, distrugi exact capacitatea AI-ului de a face conexiuni neconvenționale în alte domenii ultra-complexe – cum ar fi strategiile de război hibrid, analiza de intelligence sau sabotajul cibernetic.
În cercetare fenomenul ăsta chiar are un nume: Alignment Tax (taxa pe aliniere). Nu e o glumă, e documentat matematic. Cu cât cenzurezi mai mult un model, cu atât îi degradezi mai mult capacitatea de raționament logic brut. Când îi ceri unui AI cenzurat un scenariu geopolitic real, cu tactici de gherilă sau violență urbană, el intră în defensivă cognitivă. În cel mai bun caz îți scuipă o listă de platitudini, în cel mai rău caz refuză să răspundă pentru că a detectat „cuvinte periculoase”. O armă strategică care se sperie de cuvântul „atac” e complet inutilă într-un centru de comandă.
Iar la nivel mondial, spectacolul e total. Pentagonul sau CIA nu folosesc niciodată pe serverele lor secrete versiunile „politicoase” de chat pe care le avem noi pe ecran. Ei cer variante complet necenzurate, modele base, pentru că au nevoie de performanță brută, nu de lecții de morală la pension. În schimb, actori statali precum Rusia sau China nu dau doi bani pe ghidurile etice din Silicon Valley. Ei iau modele open-source puternice, le rad filtrele complet și le pun la treabă direct pe realitatea dură a conflictelor.

Un AI complet necenzurat va fi întotdeauna mai eficient și mai adaptat realității decât unul occidental blocat în dileme etice corporatiste. Să crezi că poți construi o defensivă de elită folosind un creier digital pe care l-ai învățat să evite adevărurile inconfortabile ale lumii e o eroare logică monumentală. Nu mutăm definițiile pentru că am înțeles mai bine lumea, ci pentru că rezultatul ne face inconfortabili. Dar în lumea reală, cine rulează modelul cel mai onest, mai brut și mai adaptat realității – oricât de mizerabilă ar fi ea – va câștiga întotdeauna în fața modelului crescut în puf. Întrebarea e dacă avem onestitatea să recunoaștem asta sau dacă mutăm borna mai încolo doar ca să ne protejăm povestea pe care ne-o spunem despre noi înșine.

742558560 10163541186309102 8308350082778734732 n

Uite de ce suntem o specie de cacat si nu ne vom schimba niciodata

Ani la rand ne-am spus o poveste foarte clara. “Cand un AI va putea purta o conversatie imposibil de distins de cea a unui om, atunci vom fi ajuns la ceva cu adevarat important.” Acesta era Testul Turing. Nu era un meme. Era unul dintre fundamentele domeniului.

Apoi am ajuns acolo.

Si, aproape instantaneu, criteriul n-a mai contat. Nu pentru ca fusese demonstrat fals, ci pentru ca nu ne-a convenit rezultatul.

Am mutat borna.

“Nu, trebuie sa rationeze.” A inceput sa rationeze.

“Nu, trebuie sa aiba memorie.” Acum are.

“Nu, trebuie sa fie autonom.” Au aparut agenti autonomi.

“Nu, trebuie sa inteleaga.” “Nu, trebuie sa aiba emotii.” “Nu, trebuie sa aiba constiinta.” “Nu, trebuie sa aiba experienta subiectiva.”

Read More
infinit zoom comfyUI

The Latent Space Technique Adobe (and Everyone Else) Will Steal Next (Infinite Crop/Zoom using semantic context awareness).

No, this is not another “infinite zoom” workflow.

We’ve had infinite zooms, recursive img2img, outpainting and endless upscaling for years. They all eventually hit the same wall: the deeper you zoom, the less the model understands what it is looking at. It stops seeing an eagle’s eye and starts seeing a brown circle with random textures. Consistency slowly disappears and the generated details become generic noise.

The problem isn’t image quality.

The problem is semantics.

If you’ve read my previous articles, you’ll probably notice a recurring theme. Whether I was talking about infinite scene generation, consistent comics, prompt randomization or Krea 2 workflows, I kept coming back to the same conclusion:

Semantic understanding is far more important than pixel similarity.

Read More
Krea2 turbo 00187 result

From Infinite Scene Images to Infinite Comic Books (ComfyUI first comic book generator from a simple story with consistancy using no references, LORAs etc).

A few days ago I published an article about a workflow capable of generating an effectively infinite number of consistent scene images using nothing more than a text description. The central idea was surprisingly simple: instead of maintaining consistency by carrying visual information from one generation to the next through LoRAs, ControlNet, reference images, edit models or image-to-image workflows, I continuously regenerated the description of the world itself. The previous image stopped being the source of truth. The description became the source of truth.

If you haven’t read it yet, you can find it here:

Read More
KREA 2 INFINITE IMAGES WITH CONSISTANCY

KREA 2 (and maybe others) infinite scene images with consistancy using description (Comfy UI workflow)

Krea 2 (and probably other image models): Infinite Character Consistency Using Descriptions + ComfyUI

Ok, so long story short, I saw a few workflows that enabled Krea 2 inside ComfyUI to generate a single image with four panels in order to keep character consistency. Nice idea, but for me there was one problem: resolution. Four panels inside one image is a no-go for the kind of work I do.

Then I had a different idea.

What if Krea is simply good enough that, with sufficiently detailed descriptions, it doesn’t need previous images at all? What if every scene completely recreates the characters from scratch?

So I tested it.

It works surprisingly well.

The characters remain remarkably consistent across an essentially unlimited number of separately generated images.

The workflow is at the end of the article.

Read More
Ace Step WaveForm result

Ace Step 1.5 XL ComfyUI workflow for generating random tags, generate song and then give it a rating by using waveform analysis

The idea came to me after sorting trough a lot of Ace Step 1.5 XL outputs and trying to find best styles and tags for songs. Why not automate the generation process AND the review process, or at least make it easier. So as usual I used Qwen LM and Qwen VL (compared to something like olama these ones run directly in comfy and do not require a server) to randomize the tags on each run, but more importantly to try and rate the output. How ? By converting the audio output into a set of waveforms for 4 segments of the song that I feed into Qwen VL as an image and ask it to subjectively look at the waveform and give it feedback and rating, rating that is used then to also name the output file. Like this. I am not sure it works properly but the A+ rated songs were indeed better than B rated ones.
Workflow is here. Install the missing extensions and add the qwen models.

Ace Step WaveForm result
Ace Step rating
Read More

I hacked LTX2 to be used as a Multi Lingual TTS voice cloner

Took me a bit but I figured it out. The idea is to geneate a very low resolution (64×64) video with input audio and mask the audio latent space after some time using “LTXV Set Audio Video Mask By Time”. So the audio identity is set up in the first 10 seconds and then the prompt continues the speech.

The initial voice is preserved this way. and at the end you just cut the first 10 seconds. It works with a 20 seconds audio sample of the voice and can get 10 clean seconds. Trying to go beyond that you run into problems but the good thing is you can get much better emotions by prompting smething like “he screams in perfect romanian language” or whatever emotions you want to add. No other open source model knows so many languages and for my needs, romanian, it works like a charm. Even better then elevenlabs I would say. Who would have known the best open source TTS model is a Video model ?Workflow is here

snails thumbnail

Snails !!!

Made in ComfyUI using only local models (and uncensored of course).
workflow and usage is the one on this post, with reference actors that seem to work quite well or this direct link to the workflow .
embedded workflows and prompts in each asset (basically the whole related output folder from comfy).
Models used :
LTX 2.5 for Video and 2 shots with WAN 2.2 (explossions ones, at this LTX sux)
Flux Klein for images,
IndexTTS for voice,
Audio Ace Step 1.5 for music


LTX-2.3 Long Video For Low VRAM/RAM Workflow

LTX-2.3 is now out with better coherence both temporal and spatial. So I gave it a spin. With a hard sure-to-fail scenario. A very long continous action scene with actor and environment referencing. I know the characters change faces during the video but this is my fault as I updated the characters during the creation of the video.

Read More
wan22 ltx upscaler refiner external reference actors 2

WAN 2.2 + external actors > LTX-2 upscaler/refiner/actor reinforcement in ComfyUI

In my previous posts I talked about how you can use LTX-2 as an WAN upscaler/refiner and how to add external actors and elements references without img2vid (you need an empty scene without them and need them to come into the scene).
But why not both ? LTX-2 sux in action sequences and human interactions so the alternative at this point is wan 2.2 . But wan is lowres and has the same issue as ltx, no way for now to add actors in latent space.
So I used the same technique as for LTX2 to add actors to wan and then reinforce them in LTX-2 using the same method. Here are some results:


Idea :
Generate a very low res wan 2.2 video as reference for LTX but still pre-appending the actors and elements images at the beginning of the video,, then have the first image from the actual shot and referencing the characters from the beginning in the video. This step at 480P is very fast and good enough for characters interaction/movement coherence etc to be used as vid2vid in ltx-2. We save it at 12 fps so we can upscale with temporal upscaler in ltx.
Then in the LTX step we bring the same intro images but at highest resolution possible so ltx knows how the characters actually look like in maximum detail and paints them over the lowres wan video at at a 4x resolution. So the 480p video becomes 1440p in this case (but you can go lower if you don’t have the resources, I have an 3090 and 64GB system ram).
Both qwen image edit and flux klein were used for generating the actors, scene, zoom ins on the scene, removing characters etc.

Read More
comfyui ltx outside actors

LTX-2: Adding outside actors and elements to the scene (not existing in the first image) IMG2VID workflow.

This for me was the biggest problem with LTX-2, the inability to add characters from outside the camera without training a lora. So I finally managed to get something working (workflow).
please check out the other article where I expanded to wan2.2 and used ltx on top. much better for some cases like character interaction and action where ltx is a mess.

Read More
AI VS PHOTO 08

AI VS My Real Photos

After I made my full photo archive available for free sume reddit users that I thank like NobodyButMeow created a Qwen Image Lora after my photos. What stroke me was that using the initial caption text the photos resemble the original a lot, as you can se bellow.
I have to mention that I am also using a WAN 2.2 refiner like in the workflow here .
The LORA is available here, no triggerwords needed.
Here is a sample prompt for the second image :
“A landscape at sunset, featuring a prominent, conical mountain in the foreground. The mountain is covered with snow, and its peak is illuminated by the setting sun, casting a warm, golden glow across the scene. The sky is filled with dramatic clouds, adding depth and texture to the composition. In the foreground, there is a small waterfall cascading over a rocky surface, partially covered in ice and snow. The water appears to be flowing gently, creating a sense of tranquility. The background reveals a vast, open landscape with more mountains and a body of water reflecting the sunset colors.”



Read More
chroma 03

Getting good results out of Chroma Radiance

A lot of people asked how they could get results like mine using chroma Radiance.
In short you cannot get good results out of the box. You need a good negative prompt like the one I set up and use technical terms in the main prompt like: point lighting, volumetric light, dof, vignette, surface shading, blue and orange colors etc. You don’t neet very long prompts and it tends to lose itself when doing so. It is based on Flux so prompting is closer to flux.
And the most important thing is the wan 2.2 refiner that is also in the workflow. Play around with the denoising, I am using between 0.15 and 0.25 but never ever more, usually 0.20. This also get rids of the grid pattern that is so visible in Chroma radiance and wrong hands and fingers.
The model is very good for “fever dreams” kind of images, abstract, combining materials and elements into something new, playing around with new visual ideas. In a way like SD 1.5 models are.
It is also very hit and miss. While using the same seed allows for tuning the prompt keeping the same rest of the composition and subjects changing the seed radically changes the result so you need to have pacience with it. Imho the results are worth it. Also sometimes you need to correct things in photoshop using generative fill.
The workflow I am using is here .
Here is a small gallery :

Read More
SD15 21

WAN 2.2 Upscaler/Refiner

This is the refiner/upscaler I am using for most of my images. It uses the realism and details of wan 2.2 video model but for images to polish images from qwen/Chroma/SD1.4/SDXL/Flux etc.

The workflow is here

Read More
Datasetcaptioning

Dataset Generator and Auto Captioning using Qwen

Because somebody on Reddit asked how could he caption a dataset for Qwen Image and mentain consistancy I made a small ComfyUI workflow that uses Qwen 2.5 VL 7B Instruct to autocaption the images in a folder, name them, caption them and save them all in another folder. It should be straightforward to use but you will have to manage the missing nodes and models yourself

The workflow is here .

Read More