Concept
The selection of this topic stems from an interest in the aesthetics of vintage botanical illustration—a visual language that merges scientific precision, decorative artistry, and historical print culture. Unlike modern digital graphics, antique botanical plates possess a stable set of defining characteristics: a central composition featuring a single plant, a vertical format, aged paper texture, an ornamental border, and a soft, muted color palette.
The project’s objective is to train the generative model Stable Diffusion XL to reproduce the visual style of vintage botanical illustrations and subsequently use it to create new, non-existent floral images within that same artistic language. A critical component of this project was assessing the model’s capacity not just to mimic specific images, but rather to internalize consistent stylistic traits: composition, color palette, texture, and the overall character of antique printed illustration.


Examples of images used to train the generative model
The chosen theme is conducive to training the model, even with a small dataset, because all images share high visual coherence. They are unified by a similar compositional scheme, a uniform presentation of the object, a framing structure, historic print texture, and a retro palette. Because of this, even a limited set of examples allows the model to grasp the overall style and reproduce it in new generations.
Dataset
A compact dataset of 12 vintage botanical cards was compiled to train the model. The visual source material consisted of historical flower images from The Metropolitan Museum of Art collection.


Examples of images used to train the generative model
During the dataset preparation phase, the images were standardized to a unified format. The lower portion of the captioned cards was partially removed to prevent the model from generating random text or pseudo-characters. Subsequently, the images were centered and cropped to a square 1:1 aspect ratio at a resolution of 512 × 512 pixels.

When selecting images, special attention was paid not to the botanical accuracy of specific species, but rather to the repeatability of the visual language. Specifically: the vertical presentation of the object, a limited color palette, a decorative border, a vintage paper texture, and the isolated placement of the plant in the center of the image.
The dataset included cards where the main stylistic characteristics were clearly legible.


Example images used to train the generative model
Model Training Process
The work was performed in the Google Colab environment utilizing a GPU. Stable Diffusion XL 1.0 was selected as the base model, and fine-tuning was carried out using the DreamBooth + LoRA method. This approach allows us to adapt the model to a specific visual style using a relatively small dataset, rather than retraining the entire model.


Image examples used to train the generative model
In the first phase, the source images were uploaded to Colab on Wednesday, after which preliminary processing was performed: archive unpacking, cropping, removal of the bottom sections with signatures, and standardization of the files to a unified resolution of 512 × 512 pixels. The prepared images were saved in a separate folder and used as the training dataset.
The main training parameters were then set: the path to the dataset, the results saving directory, a textual description of the target style, and the number of training steps. The primary prompt used was the phrasing «a botanicardstyle vintage botanical card, ” where the unique token botanicardstyle denoted a new stylistic attribute the model was required to assimilate.
Training was performed using the script train_dreambooth_lora_sdxl.py. During the process, lightweight settings were chosen to enable training within constraints of limited time and resources: a 512-pixel resolution, a small number of steps, the use of mixed-precision computing, and memory optimizations. This allowed us to successfully obtain the LoRA weights and subsequently connect them to the base model to generate the final series of images.
Upon completing the coursework, a base Stable Diffusion XL model was loaded, to which the derived LoRA weights were connected. The model was then utilized to generate new images based on text prompts. All prompts were structured around the common template of a vintage botanical illustration, differing only in the type of fictional flower, color characteristics, and atmosphere.
Notebook Link: Open Google Colab
Resulting images and their analysis


Resulting Images
As a result of the course, a series of new images was generated in the style of vintage botanical illustrations. Although the final flowers are fictional, most of the works successfully retain the key characteristics of the training dataset: central composition, isolated object, soft retro color palette, the texture of aged paper, and characteristic decorative presentation.
Resulting Images
The model most successfully reproduces the general card format and the character of the floral object. The elongated shapes of the stems, petals, and inflorescences are particularly well-rendered, as is the overall style of an antique printed illustration. Because of this, the images are perceived as elements of a single conceptual botanical atlas, even if the flowers themselves do not exist in reality.
However, the results vary in terms of accuracy. In the most successful generations, both the flower’s structure and the overall compositional logic of the card are well-preserved. In less stable examples, individual artifacts may appear: inaccuracies in petal shape, unnatural stem bends, uneven detailing, or an arbitrary frame. Nevertheless, these deviations do not disrupt the overall stylistic unity of the series.


