Project Idea
The project idea came about quite accidentally—I stumbled upon an old craft project: a model of the solar system made from a simple cardboard box. It immediately caught my attention: the uneven planets, the basic paint job, the slightly naive proportions, and the feeling that everything was assembled by hand without trying to be «perfect.»
This object possesses a distinct atmosphere because it is perceived not as a scientific model, but as a personal interpretation of the cosmos.
I became interested in whether this feeling could be translated into a generative model and maintained when creating new images. As part of this project, I am training Stable Diffusion on photographs of this craft to then generate a series of images where the object retains its recognizability but exists under different conditions—with new lighting, composition, and environment.
Thus, I view the model not merely as a generative tool, but as a means of reinterpreting the original object—examining how well it can capture its essence and continue its visual logic beyond the original form.
The resulting images can subsequently be used as a basis for illustrations or visual series, where the children’s interpretation of the cosmos transforms into an independent artistic motif.
For training the model, I independently compiled a dataset of craft photos. The object was shot from various angles to capture its form and volume, and also under different lighting conditions—ranging from soft to high-contrast.
All images were converted to a square format. The final dataset comprises 20 photos that vary in shooting angle and light, making it more diverse and suitable for training.
Application of a generative model
The following tools were utilized during the project: — Stable Diffusion for training the generative model; — Google Colab for code execution; — ChatGPT for prompt writing.
Training a Generative Model
1) Environment Setup and Dependency Installation
In the initial phase, the working environment was configured in Google Colab, and the necessary libraries for working with Stable Diffusion were installed.
2) Project Configuration
Next, I set the main project parameters, including the path to the dataset and the textual identifier of the object used during training.
3) Verify the dataset and image preview
At this stage, the data folder is checked, suitable images are selected, and their previews are displayed.
4) Additional Preprocessing
At this stage, additional data processing is performed: images are converted to RGB format and are normalized to a square aspect ratio with centering.
5) Creating Signatures / Metadata
In this section, a metadata file is generated for use during the learning process.
6) Setting Up Accelerate
At this stage, the Accelerate configuration is prepared to run the training script.
7) Model Training (LoRA)
At this stage, the environment is tested, the team is assembled, and the model training is initiated.
8) Check Saved Files
After completing the training, a list of files created in the output directory is displayed.
9) LoRA Integration
The model loads, and a trained LoRA is applied to generate images.
Final Images
After completing the training, I proceeded to image generation using the trained model. For this, I utilized my own text prompts, specifying the subject and setting various environmental and lighting conditions.
All prompts were built around a unique object identifier to ensure the model accurately reproduced the trained form and retained its visual characteristics across different scenes.
prompt: «educational poster style image of handmade solar system model, clean background, vibrant colors» The initial test involved generating an image using a trained model, where the object was maintained under conditions closely approximating the original against a completely white background. The result was immediately impressive: the neural network accurately conveyed the form of the craft, the placement of the planets, and the overall composition.
Stable Diffusion — «Solar System, ” 2026
prompt: «minimalist photo of handmade solar system model on a white pedestal, gallery lighting»
In the next generation, the object retains its recognizable structure, but it is interpreted slightly differently by the model. It is clear that the shape and placement of the elements remain close to the original, yet the labels and fine details begin to distort, highlighting the nuances of the neural network’s operation.
Crucially, the sense of handmade texture and paint application is well preserved, which allows the image to still look like a physical craft rather than a completely digital object.
