Original size 960x1280

Apartment Totem

PROTECT STATUS: not protected
Longread translated automatically
The project is taking part in the competition

A potted plant usually has no role—just a spot on the windowsill. I trained Stable Diffusion XL myself to keep this exact specimen as a character. From there, the totem is placed into other settings: a bus stop, a roof, a museum, a forest.

The assignment was to train a generative model on my own style or a specific subject. I chose a subject, not a painter’s style. There are three reasons for this. The entire dataset is mine: 14 square photographs of a plant on a table, with no stock imagery or other illustrations. Consistency is checked simply: can the same flower be recognized? The point of the project is not to imitate a canon, but to pull a mundane object out of the background and give it a biography. Working token: TOK totem. This is a rare name so the model won’t draw «any flower.»

Dataset

14 frames, 1:1 aspect ratio, original photography. Angles: top-down, side profile, close-ups on the leaves and pot, under table light. The shooting objective was very narrow: to capture the silhouette, the volume of the foliage, and the character of the pot, without changing the «actor.» The background was almost uniformly the same—a table. This was a risk: the model might associate the tablecloth with the plant. Therefore, in the generation process, I intentionally placed the totem in different environments.

Learning

Base: Stable Diffusion XL 1.0Method: DreamBooth + LoRAEnvironment: Google Colab, T4 GPUCourse laptop, tok folder instead of the tutorial example.
Parameters:instance_prompt: a photo of TOK totemResolution 512, 200 steps, learning rate 1e-4, 8-bit Adam, fp16.BLIP was used to caption the frames; all lines begin with a token.
200 steps—a short run: the Colab environment initially crashed due to a conflict between the diffusers / torchao libraries, and a full cycle of 500–600 steps would not have fit within the deadline. For small anchor datasets, the current steps are usually sufficient. The text encoder was not fine-tuned.

big
Original size 1919x944

Series

My generations can be found at this Yandex Disk link:
https://disk.yandex.ru/d/BSoQWfJg4l7k6g

The prompt always begins with a photo of the TOK totem. Only the scene changes: a table at dawn, a museum display case, the roof of a panel building, a night bus stop, a drafting table, a forest, a studio portrait, a windowsill at night.
Negative prompt: blurry, deformed, watermark, text.

No additional networks (ControlNet, IP-Adapter) are used. If the frame included my table, I used a smaller lora_scale.

Analysis

What held together: the overall silhouette of the plant, the character of the foliage, the feeling of a potted plant rather than a field bouquet. The totem reads best in calm shots (studio, desk, windowsill). What broke down: the fine structure of the leaves is blurry; sometimes an «extra» flower appears; in urban scenes, the model tends to affix a piece of a table or the domestic lighting is too subdued. This is the price of a short LoRA and a homogeneous background in the dataset.

Connection to the Idea: The project isn’t about producing a perfect catalog render. The test is simple—can you recognize the same occupant at the desk if you move them to a bus stop or into the forest? Where recognition occurs, the concept works. Where it remains a «mere plant, ” the boundary of the small sample is visible. Series Variations: Everyday shots are closer to the learning process and more stable; night and city test a new light; the museum and monument change the object’s status; the forest and dark room serve as a stress test without the comfort of apartment cues.

Conclusion

LoRA on 14 frames doesn’t create a new artist. It creates an anchor. This anchor was enough for a plant to stop being a background element and become a character.

Application of Generative Models

The core model is a fine-tuned Stable Diffusion XL (DreamBooth + LoRA) developed by me. Base model: https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0 Training notebook: https://colab.research.google.com/github/shwars/ai-for-creatives/blob/main/6-DiffusionTraining/SDXL_DreamBooth_LoRA_Colab.ipynb

Dataset captions were generated using BLIP (Salesforce/blip-image-captioning-base), which is part of the provided notebook.

The structure and draft text were generated using Grok (xAI, Grok 4.6). The phrasing has been manually shortened and revised to align with the plant and the actual shots. The final images in the series exclusively use the trained LoRA.

No other generative models were used for the final images.

Apartment Totem
Project created at 26.09.2026