Exactly ten years ago, I decided to make a virtual reality application. I did not know Unity. I did not know C#. I learned enough of both in approximately two weeks and built Window to the World—to the best of my knowledge, the first Romanian VR application released on a public app store.
The application is still listed on the Meta Store:
https://www.meta.com/experiences/window-to-the-world/977682772321493/
But the project did not begin with VR. It began years earlier with photography and a simple frustration: a photograph could preserve light, colour, composition and detail, but it could not fully preserve the physical sensation of being there.

Photography Without a Drone

In 2013, I started experimenting with ways of giving my photographs depth and movement. I did not own a drone, and consumer drones were not yet ordinary photographic tools. If I wanted a camera movement through one of my landscapes, I could not simply return to Iceland, France or Bulgaria and fly through the scene. So I tried to reconstruct that movement from the photograph itself.
I separated the images into elements, recreated the areas hidden behind foreground objects, built approximate geometry in 3ds Max, projected the original photographs onto it and added reflections, displacement, animated water and camera movement. It was not the usual two-dimensional parallax effect made from a few flat layers. The photographs became limited three-dimensional environments. They still worked only within a restricted range of movement, but inside that range the camera existed in something closer to an actual space.
I documented the process here:
https://aurelm.com/2013/12/15/fernweh-3d-parallax-photography-tutorial/


The original project was a video: a sequence of photographs transformed into small moving worlds. But once the scenes existed in three dimensions, rendering them as conventional video began to feel unnecessarily restrictive. The viewer could still only follow the camera movement I had chosen. The natural question was: why not allow the viewer to control it?
The next version moved the project into WebGL. Instead of watching a predefined animation, the viewer could move the mouse and change the perspective in real time. I also added genuine interaction with the scene: water reacted with ripples, particles moved through the environment and different elements responded to the viewer.
I again spent roughly two intense weeks teaching myself what I needed—this time JavaScript and Three.js—and transformed the parallax project into an interactive real-time experience. It stopped being only photography and became an attempt to find the border where photography, three-dimensional graphics and interactive media could meet.
But the viewer was still looking through a monitor. The scene responded to the person, yet remained trapped inside a rectangle. VR was the obvious final step.

The First VR A-Ha Moment

My interest in virtual reality had begun even earlier. In 2012, I attended SIGGRAPH in Los Angeles. Among the experimental technologies exhibited there was a rough virtual reality prototype.
It was a large black device connected to a computer by a ridiculous number of cables. The resolution was terrible and the response time was far from ideal. Yet it took only a few seconds for my senses to accept the illusion.
The experience placed me inside the cockpit of a stationary racing car. My rational mind knew that nothing around me was real, but that knowledge made little difference to the rest of my brain. My immediate reaction was to reach for the virtual steering wheel. Then came a strange moment of cognitive dissonance when I saw the virtual hands pass through it. Intellectually, I understood what was happening. At a subconscious level, my body had already accepted the space as real.
I left the chair with the powerful feeling that I had just experienced a fragment of what was coming.
That prototype would eventually evolve into the first Oculus development kit. Later, a DK1 reached the Gameloft studio where I worked, giving me several dozen hours to experiment with it. The image quality and latency were still poor, but the direction was unmistakable.
Two years after SIGGRAPH, in 2014, I bought an Oculus DK2 for my home. I wrote about my first week with it here:
https://aurelm.com/2014/10/26/dupa-o-saptamana-de-oculus/
The DK2 was primitive by current standards, but it was an enormous improvement. It had better resolution, contrast, latency and positional tracking. The gaps between pixels remained obvious, the software was difficult to configure and many applications required bizarre rituals just to start correctly. None of this mattered very much.
Alien: Isolation made my entire body react as if the danger were real, despite my mind constantly repeating that it was only a game. Virtual cinemas made it possible to sit inside an enormous theatre at home. Experiences such as Senza Peso and Sightline suggested entirely new forms of cinema, narrative and psychological interaction.
I had a similar reaction when I tried Google Glass. It was clearly not the final form of augmented reality, just as those Oculus prototypes were not the final form of VR, but it demonstrated the direction.
I became convinced that augmented reality would eventually replace the smartphone. I still believe that completely.
It will happen once a technological and cognitive threshold is crossed. The devices must become light, socially acceptable, visually convincing and useful enough that wearing them requires less effort than repeatedly taking a rectangle out of your pocket. Once that threshold is reached, the transition may happen much faster than most people expect.
The smartphone will not necessarily disappear. It will simply become less central as the interface migrates from an object in our hands into the space around us.

Building Window to the World

When consumer VR finally became accessible through Samsung Gear VR, I realised that the old parallax photography project could become what it had always been trying to become.
At the time, there was no mature ecosystem of affordable standalone VR headsets. Gear VR was the first practical consumer platform with a real application store. You inserted a compatible Samsung phone into the headset, placed it on your face and entered VR.
I already understood photography, Photoshop, 3D modelling, texturing, shaders, real-time graphics, compositing and optimisation. What I did not know was Unity or C#.
So I learned them—not completely or academically, but exactly as much as I needed to solve the problems standing between the idea and a working application.
In approximately two weeks, I transformed the project into a VR experience. The viewer could enter a collection of reconstructed photographic moments, look around them, perceive depth and experience moving water, reflections, particles and other environmental elements.
My ambition was simple to describe and difficult to achieve: to express photography as closely as possible to the actual moment of capture, approaching the point where photography and reality begin to converge.
The evolution now seems almost inevitable: first a photograph, then a video simulating movement inside it, then a WebGL experience allowing the viewer to control and interact with it, and finally a VR application in which the viewer was placed inside it.
That is why I called it Window to the World.
Before its official Oculus Store release, the application was distributed through SideLoadVR and received an extraordinarily enthusiastic review from Gear VR News:
https://gearvrnews.wordpress.com/2016/03/23/window-to-the-world/
The reviewer described the images as staggering and the illusion as remarkably convincing. He compared it to opening an entire wall of your home towards another place on the planet and said it had the polish one might expect from a major studio, not from one person who had just taught himself the required programming.
Most importantly, he understood what I was trying to create. He did not describe it merely as a gallery or a parallax experiment. He described it as a window into another place and, ultimately, as a window into the future.

Two Weeks Does Not Really Mean Two Weeks

Saying that I built the application in two weeks is true, but also misleading. The Unity and C# part took approximately two weeks. The application itself was the result of years spent learning photography, Photoshop, 3D modelling, texturing, real-time graphics, compositing, optimisation and visual perception.
This is often how apparently sudden technical achievements happen. A person learns a tool in two weeks, but that tool attaches itself to decades of previous knowledge. The final step is fast because the structure underneath it was built slowly.
I did not become a conventional software engineer during those two weeks. I learned the amount of programming required to connect abilities I already possessed. Code was not the destination. It was the missing bridge.
The most interesting work frequently happens between disciplines, where nobody possesses all the officially required qualifications and therefore has to invent a path forward.

A Similar A-Ha Moment

At the beginning of 2022, I experienced an a-ha moment similar to the one I had experienced with early VR. This time it happened on my own computer, when I discovered Disco Diffusion and generative AI.
The images were slow, unpredictable and often broken. Faces were distorted, details dissolved into visual noise and prompts behaved less like commands than negotiations with an alien visual intelligence. But again, the imperfections were less important than the direction.
I immediately understood that this was not merely another image-processing tool. It was the beginning of a new interface between human intention and digital creation.
From that moment, I began writing intensively about how AI and generative models would change the world. I practically invaded my friends’ feeds with posts about it. Most people were sceptical. Some saw the images as toys, random curiosities or another temporary technological trend.
To me, the consequences were already obvious.
I later became a beta tester for Midjourney, DALL·E and, most importantly, Stable Diffusion. Stable Diffusion represented something fundamentally different. It could run locally, be modified and trained by individual users, and was not entirely controlled through a closed company interface.
You were no longer merely asking a remote service to generate an image. You could possess the model, run it on your hardware, change its behaviour and train it on new people and concepts.
During that year, I trained Stable Diffusion on my own face, my father’s face and several other people, publishing galleries of the results on Facebook long before custom AI portraits became an ordinary product.
Photography naturally became one of my first AI playgrounds. In the summer of 2022, while Stable Diffusion was still in beta, I generated synthetic landscapes that reproduced much of the visual grammar of landscape photography: atmospheric depth, dramatic light, mountains, glaciers, reflections, clouds and carefully constructed foregrounds.
The results caused considerable debate after being published by PetaPixel:
https://petapixel.com/2022/08/16/these-are-not-photos-beautiful-landscapes-created-by-new-ai/
The article touched an uncomfortable question: if an image can produce the visual and emotional effect of landscape photography without depicting a real landscape, what exactly is the viewer consuming? The journey, the physical location, the authenticity of the capture, the technical process—or simply the emotional force of the final image?
This was not a new Photoshop filter. It was a change in who could create visual worlds, how those worlds could be produced and who controlled the tools.

Presenting the Generative Ecosystem at Dev.Play

Later in 2022, I presented and discussed these emerging image-generation methods on stage at Dev.Play:
https://www.youtube.com/watch?v=2frapV4JZz0
The presentation was not limited to typing a prompt and receiving an image. I presented the broader creative ecosystem growing around Stable Diffusion and Automatic1111.
Compared with the closed generators available at the time, Automatic1111 offered something much more important than text-to-image. It offered an open environment containing img2img, inpainting, outpainting, prompt weighting, upscaling, scripts, extensions and an expanding collection of community-built tools.
Img2img allowed a photograph, sketch or render to become the starting point for a generation. Inpainting allowed one region to be regenerated while preserving the rest. Outpainting extended an image beyond its original frame. Stable Diffusion was no longer merely producing isolated images from sentences. It was becoming a complete generative ecosystem.
That distinction mattered enormously to me.
A closed generator gives you results. An open ecosystem gives you a medium.
Towards the end of the year, ChatGPT appeared. The scale of the change became visible to almost everyone.
And the rest is history.

From Automatic1111 to ComfyUI

Much of what I saw developing around Stable Diffusion and Automatic1111 later materialised, for me, in ComfyUI.
Today, ComfyUI is my creative platform in probably 99% of cases involving generative images and video. Its node-based structure allows me to treat generative models not as isolated services, but as components inside larger systems.
I use it for image and video generation, but also for experiments that combine large language models with image and video models. An LLM can analyse an image, generate or rewrite prompts, preserve context between generations, choose parameters or modify the next stage of a workflow.
A prompt does not need to be written once by a human and sent directly to one model. It can be generated, analysed, expanded, corrected and adapted by other models. An image can be examined by a vision-language model, transformed by an image model, interpreted again, used to produce video and then passed into another stage.
ComfyUI is therefore more than my preferred generation interface. It is an extremely useful environment for prototyping technical and creative ideas.
A workflow can begin as an experimental node graph, perhaps supported by one custom node written for a very specific operation. Once the idea works, it can be transformed into a dedicated application with its own interface, controls and workflow, hiding the technical complexity from the final user.
A recent example is DYSTALGIA: Infinite, Context-Aware Generative Zoom:
https://aurelm.com/2026/07/18/dystalgia-infinite-context-aware-generative-zoom/
The idea began with a question: can a generative model zoom into one part of an image without immediately forgetting the world around it?
Traditional upscalers enlarge the information already present. This technique instead analyses both the selected crop and the complete original image. The original image acts as semantic context, helping the system understand what the crop represents, how it relates to the larger composition and what kind of detail would make sense inside it.
The result becomes a new image into which the user can zoom again. Each level can become the starting point for another crop and another generation, while the original scene remains part of the context. It is an interactive exploration through latent space rather than a predefined infinite-zoom animation.
The concept first existed as a ComfyUI workflow and a custom crop node. Once the workflow proved that the idea worked, I used Codex to transform it into an application with its own interface in roughly half a day.
DYSTALGIA therefore demonstrates another important role for ComfyUI. It is not only a platform for producing images. It is a technical sketchbook in which a generative application can be designed, tested and refined before becoming a standalone product.
The same workflow logic can later sit behind a much simpler interface. The person using the final application does not need to understand nodes, model loaders, conditioning, samplers or latent representations. They only need to load an image, select an area and continue travelling into it.
This is where generative AI becomes more than a collection of tools. It becomes a programmable creative environment and, increasingly, a foundation for building entirely new applications.

Creativity Must Be Free

I believe profoundly in open-source software and locally running models. I believe this is where the most important forms of genuine creativity will survive and develop.
Real creativity does not accept corporate censorship as its final boundary. It does not ask a platform which thoughts, images, narratives or experiments are permitted by its current policies.
Creativity can accept technical limitations. Every medium has them. A painter works within the properties of paint. A photographer works within light, optics and time. A generative artist works within the architecture, training and computational limits of a model.
Those are limitations of the medium. Corporate restrictions on thought and ideas are something else.
Large companies increasingly try to determine not only which outputs are illegal or directly harmful, but which subjects, visual ideas and forms of expression users are permitted to explore. Those boundaries can change overnight according to opaque policies designed primarily to protect a company rather than enable an artist.
That is incompatible with real creative freedom.
This does not mean that every generated result is valuable, moral or worth publishing. Creative freedom does not remove personal responsibility. It means that the final decision belongs to the creator, not to a remote corporation silently modifying the boundaries of imagination.
A locally running model can be studied, adapted, trained, combined with other systems and used privately. It does not disappear because a company changes direction. It does not require permission every time the user explores an uncomfortable idea.
Open source does not guarantee creativity, but it preserves the possibility of it.
Real creativity is free—in thought and ideas, even when constrained by the technical limits of its tools.

Ten Years Later, I No Longer Search for Applications

Over the last few days, another transition has become visible to me. I have started making applications almost continuously.
I no longer begin by searching for an application that approximately solves a problem. In many cases, it is easier to build exactly what I need than to search through dozens of products, install them, discover their limitations and adapt myself to somebody else’s assumptions.
This is not because I suddenly became an expert software engineer. It is because AI has changed what it means to create software.
Natural language has become a practical interface between an idea and working code. My experience in graphics, interaction, photography and technical art still matters enormously, just as it did when I learned enough Unity and C# to create Window to the World. But the distance between understanding what I want and producing a functional application has collapsed.
In only a few days, I have created multiple Android and Windows applications, alongside Python-based tools that should be adaptable to several operating systems.
One is Me, the Puzzle:
https://aurelm.com/2026/07/20/me-the-puzzle-reassembling-a-fragmented-reflection/
The application uses the phone’s live camera feed to turn the person holding it into a shattered, moving puzzle. Instead of rebuilding an abstract image, you reconstruct your own reflection. Each fragment carries part of the live camera feed and reacts through depth, movement, reflections, edge lighting, distortion and occasional digital glitches.
Even connected pieces can still be moved, so becoming whole is not represented as a permanent state, but as something continuously assembled, disturbed and reconstructed.
It began as a personal metaphor for memory, mood and identity, but it became a game people genuinely wanted to try. Friends installed the APK directly despite the warnings associated with installing an application outside Google Play. The response was much stronger than I expected.
I also built an underwater camera application specifically for my Xiaomi phone. I wanted a replacement for my late Olympus TG-5 without buying another dedicated underwater camera.
A touchscreen is unreliable when the phone is sealed inside a cheap underwater case, so I designed the entire application around the physical volume buttons. Different presses and combinations control photography, video recording, zoom and the other necessary camera functions. It now needs only an inexpensive waterproof case that does not compromise the optics.
I did not need an application designed for every phone and every possible user. I needed one designed for my phone and the precise way I intended to use it.
So I made it.
The application through which I am writing this article is another example. It is my own WordPress writing and publishing environment.
I can write the article, insert photographs directly and have their files automatically renamed according to the title and context. The application uploads and places them without the usual interruptions. At the end, it can publish either the Romanian or English article together with its automatically translated counterpart.
It includes multiple AI functions powered by a local model such as Gemma running through LM Studio. The model can correct, rewrite, develop, translate and restructure the text without sending the document to a remote AI company.
I have also built a more general word processor—something closer to Word, but designed around local AI from the beginning rather than having AI added later as a secondary feature.
These are not polished products created by large software companies. They are personal applications shaped around immediate needs. Some may remain personal, some will be shared with friends and some may eventually be sold.
But the important part is the pattern.

Personal Software Will Become Normal

For most of computing history, people adapted themselves to software. A company decided which features mattered, how the interface should work and what type of user the product was designed for.
Programming was the alternative, but it required years of specialised knowledge. For most people, the choice was simple: use what exists or go without it.
That choice is disappearing.
People will increasingly create their own applications for their own needs. They will describe a problem, work with an AI system, test the result and gradually shape the software around themselves.
Many applications will be extremely specific: an underwater camera for one phone and one case, a publishing environment for one particular website, a tool for one photographer’s workflow or a small game created for a group of friends.
Traditional software development would rarely justify building products for audiences this small. AI-assisted development changes the economics because an application no longer needs millions of users to justify its existence.
Sometimes one user is enough.
People will exchange applications with friends, families and communities. They will modify one another’s tools and create specialised versions. Some will release them for free. Some will sell them. People who would never have called themselves programmers will create personal software ecosystems around their own lives.
There are serious problems to solve: security, trust, maintenance, permissions and the difficulty of knowing whether an application actually does what its creator claims. But the success of Me, the Puzzle among people willing to install a direct APK shows that the appetite already exists.
People will accept unconventional distribution when an application offers something personal, original or useful enough.
The future may not consist only of a few million applications competing inside enormous stores. It may also contain billions of small applications, each created for one person, one family, one project or even one afternoon.

The Pattern Repeats

Looking back, the same pattern has repeated throughout this story.
In 2012, at SIGGRAPH, I experienced an early VR prototype and immediately recognised the importance of what it suggested. In 2013, I tried to recover movement and depth from my photographs. In 2014, I brought VR into my home. In 2016, I learned enough Unity and C# in two weeks to transform those experiments into Window to the World.
At the beginning of 2022, Disco Diffusion produced a similar a-ha moment. Stable Diffusion then allowed people to possess, modify and train the model itself. Automatic1111 transformed it into a growing creative ecosystem. ComfyUI turned that ecosystem into an almost infinitely configurable visual programming environment and a place where completely new applications could be prototyped.
ChatGPT allowed almost anyone to communicate with a model through ordinary language. Coding models are now doing something equally important: turning software creation from a specialised act into a conversation.
Not a completely reliable conversation. Not yet an effortless one. Experience, taste, persistence, judgement and the ability to recognise a broken result still matter enormously.
But the direction is unmistakable.
AI does not merely help us use existing tools. It increasingly allows us to make the tools we wish existed.
Ten years ago, learning Unity and C# in two weeks to build one VR application felt extraordinary.
Today, I am building applications one after another.
I no longer search for the closest available tool.
I make the exact one I need.
And soon, I suspect, so will everyone else.