Original size 1024x1536

Training a neural network on the style of Tatyana Mavrina’s work

The project is taking part in the competition

PROJECT IDEA


The main goal of this project is to train the Stable Diffusion XL model to generate images in the unique style of my childhood drawings. This will be achieved using DreamBooth and LoRA training methods.As a child, a drawing didn’t require explanation. Proportions could be off, colors could be unexpected, and objects could look nothing like their real-life prototypes. There was no need to follow composition rules, strive for realism, or adhere to a specific style.In this project, I am referencing an archive of my own childhood drawings. I am curious whether I can do more than just reproduce the old drawings—I want to continue them.

Using the generative neural network Stable Diffusion and the DreamBooth and LoRA tools, I intend to train the model to reproduce a specific aesthetic: pencil strokes, uneven paint textures, random overlay of dark colors, distorted proportions, and loose contours. The result will be a series of non-existent children’s drawings—images that never existed but could easily have appeared in my sketchbook years ago.

Below, I will detail every stage of the process: from dataset collection to the generation of the final series.

0
Original size 2480x398

Illustrations by Tatyana Mavrina

Will the model be able to adopt not only the external characteristics of the painting but also the artist’s entire visual language—first and foremost, the mood, the rhythm of the lines, the emotional expressiveness, and the overall decorative quality?Text

Original size 2480x1750

Illustrations by Tatyana Mavrina

For training the model, a dataset of 44 images was compiled, featuring a diverse body of work by Tatiana Mavrina. This collection included: landscapes, portraits, still lifes, as well as illustrations for fairy tales and literary works.

This set was necessary so the model could see style across different genres. This is important because the task wasn’t about replicating a single plot, but about transferring a general artistic language onto various scenes.

Original size 3455x1508

Original images before conversion to 512×512 format

Technical Implementation

Infrastructure Setup and Authentication:

Connecting Google Drive for data storage. Authenticating with Hugging Face Hub to access base models (notebook_login). Installing necessary libraries (diffusers, transformers, accelerate). Data Processing (Preprocessing):

Cropping: Using PIL to crop images into a square format (crop_to_square). Organization: Creating folder structures and cleaning file names. Automatic Captioning (Image Captioning):

Loading the pre-trained BLIP model (BlipForConditionalGeneration) to automatically generate textual descriptions for images. Generating a caption for each image—a step critical for the model to subsequently learn context. Configuration and Training:

Configuring the environment using the accelerate library for efficient GPU utilization. Initiating the training process (likely Dreambooth or LoRA) based on the uploaded images, utilizing training scripts from the diffusers library. Generation and Visualization (Inference):

Loading trained weights into the DiffusionPipeline. Generating new images based on text prompts. Visualizing the results using matplotlib: creating image grids (plt.imshow) to evaluate training quality.

Original size 2480x622

Code Snippet 1: Cropping images to a 512×512 format

Original size 2480x972

Code Snippet 3: Image Generation

Final Series: 10 Images and Prompts

NEURAL NETWORKS USED:

Stable Diffusion XL https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0

BLIP https://huggingface.co/Salesforce/blip-image-captioning-base

DreamBooth https://huggingface.co/papers/2208.12242

LoRA https://huggingface.co/papers/2106.09685

Training a neural network on the style of Tatyana Mavrina’s work
Project created at 24.09.2026