Training a LoRA Model for Generating Images in the Style of Soviet Diaphilms
Dataset Description
For this study, a series of images inspired by Soviet diaphilms was utilized—a distinct art form popular in the USSR from the 1930s to the 1990s. Diaphilms consisted of reels of sequential frames projected onto a screen, accompanied by narration to form an independent genre of visual storytelling.
Diaphilms are characterized by a recognizable drawn style, a soft line quality, a restricted color palette, and a unique compositional logic. The frames revolve around a central character placed within a landscape or interior, while the background and second plane are often rendered more schematically than the main subject. The flatness of the image plays a crucial role: space is constructed through patches of color and outline, rather than through perspective. A muted, warm tone, characteristic graininess, and the sense of handcrafted quality—evidence of watercolor, gouache, or ink—create a distinct atmosphere.




The project utilizes the LoRA fine-tuning method based on the Stable Diffusion XL model. The initial dataset comprises 32 images from the National Electronic Library (NEDB) and Diafilm.rf, which reflect characteristic style traits: drawn lines, narrative content, a limited palette, conventional backgrounds, and literary foundations.
Program Description
A dataset of 32 images was prepared to train the model. Before uploading, all pictures were resized to a 1024×1024 square format and saved locally. A single instance_prompt was used as the style token, describing the genre overall: «a frame from a soviet diafilm, vintage hand-drawn illustration.» Additionally, a validation prompt was set up to track the training progress. The training was conducted using the DreamBooth LoRA method based on the Stable Diffusion XL model (stabilityai/stable-diffusion-xl-base-1.0) with a VAE fix for fp16 (madebyollin/sdxl-vae-fp16-fix). The process was executed in Google Colab on a Tesla T4 GPU. To conserve video memory, gradient checkpointing, an 8-bit AdamW optimizer, and gradient accumulation with an effective batch size of 4 were employed. The final LoRA adapter had a rank of 8, which ensures a compact file size while maintaining stylistic expressiveness.
During training, the model learned the correlation between the style token and the dataset’s visual characteristics: drawn line quality, narrative composition, a restricted warm color palette, stylized background, and characteristic storytelling manner. A total of 1,500 optimization steps were executed with a cosine learning rate schedule.After training was complete, the resulting LoRA model was used to generate a new series of images based on text prompts, ensuring the style phrase-token was always preserved.
Image Series
Across all generated works, key characteristics of the diaphilm style are evident: a drawn quality, a muted warm palette, stylized backgrounds, and a handcrafted feeling. The images are united by a common aesthetic but vary in subject—ranging from winter landscapes and forest scenes to everyday vignettes.The models successfully conveyed the decorative quality, linear rhythm, and the character’s connection to the environment. The most consistent elements were the distinctive line work, the limited color scheme, and the overall «filmic» atmosphere of the frames.
