Project Concept
While traveling across America, I photographed many landscapes: roads winding between rock formations, arid plains, canyons, and mountain silhouettes. In these shots, what matters to me is not just the specific locations, but the feeling of the journey itself—the distance, the light, and the anticipation of what lies around the next bend.
In the project «The Road That Wasn’t, ” I use my personal photo archive as material to fine-tune a generative model. I am curious to see if the neural network can absorb the visual characteristics of my photographs and create new landscapes that feel like a continuation of that same journey.
Thus, an imagined itinerary emerges—its images are rooted in real experience, yet they do not document the places I actually visited.
Original photographs


«Original photos from a trip. Training dataset.»
Sixteen photographs from my American road trip were selected for this project. The collection includes desolate roads, red rock formations, canyons, natural arches, sparse vegetation, and isolated elements of infrastructure.
The shots share recurring colors and motifs: ochre earth, slate-blue skies, deep shadows, long roads, and large rock structures. Furthermore, the lighting varies—from bright sunshine to cloudy weather and sunset.
I took all the original photographs. Square versions, sized 1024 × 1024 pixels, were prepared for training without additional color correction. Each image is accompanied by a textual description of the scene.
Method
Stable Diffusion XL is used as the foundation. Fine-tuning is performed using the LoRA method: a small set of supplementary parameters is trained and subsequently attached to the base model.The project adapts a learning notebook from the «Artificial Intelligence for Creative Industries» course. The caption for the photographs uses the general phrase in alisharoad photographic style, which connects the images to the chosen photographic aesthetic.The goal of the experiment is to test which characteristics of the original archive the model can reproduce in new scenes: palette, lighting, surface texture, and compositional character.In the learning notebook, I replaced the original dataset with 16 of my own photographs from an American road trip. The image captions were prepared using Codex instead of automatic description by the BLIP model. Training was completed in 300 steps at a resolution of 512×512 pixels. The adapter weights are saved on Google Drive.
Learning Settings
For fine-tuning, the Stable Diffusion XL Base 1.0 model and the DreamBooth LoRA method were used. The dataset consisted of 16 original photographs taken during a trip through America. Training was performed in Google Colab on an NVIDIA Tesla T4 GPU at a resolution of 512 × 512 pixels. There were 300 training steps, a learning rate of 0.0001, and a LoRA rank of 4. The batch size was one image, and gradients were accumulated over four iterations before updating parameters. Seed 42 was set to ensure experimental reproducibility.
Before and After Comparison


Left is the result from base SDXL; right is the result using a trained LoRA adapter. The same text prompt and seed 100 were used.
Final Series
Results Analysis
In the final series, motifs from my photographic archive are repeated: desolate roads, red rock formations, canyons, and sparse vegetation. The images are unified by ochre tones, a muted sky, and deep shadows.
In the control pair after integrating the LoRA, the image became darker, and the color of the rocks appeared more subdued. Instead of solid rock walls, individual masses and towers emerged. These changes visually bring the result closer to my photographs. However, one control pair does not allow for this conclusion to be generalized across all prompts.
The series conveys different impressions of travel: the road evokes a sense of motion, the canyon suggests scale and enclosure, the wind turbines imply distance, and the sunset signals the conclusion of the route. This results in an imagined extension of a real journey.
The experiment has limitations. Some compositions closely resemble the training photographs, particularly scenes featuring the river’s bend, an arch, and the sunset. In two images, unsolicited human figures appeared, and certain details of the wind turbines look improbable. Therefore, the result requires visual curation and does not offer complete control over the content.
The images were generated at a resolution of 768 × 768 pixels over 25 steps, using a guidance scale of 6 and a LoRA strength of 0.8. Seeds ranging from 100 to 107 were used for the series.
Description of Generative Model Application
