Project Description
In 1856, architect Owen Jones published «The Grammar of Ornament"—a collection of one hundred and twelve lithographic plates featuring ornaments from nineteen cultures, ranging from Egyptian to Elizabethan. Jones operated under the premise that ornament functions like a language: it possesses a vocabulary of motifs and a grammar for their combination, and this grammar can be described just as syntax is described.
I was interested in what would happen if we divorced this grammar from its vocabulary. If the model absorbed not specific palmettes and meanders, but the underlying principle—flat two-color printing, mandatory regularity, counting along both the vertical and horizontal axes, and subordination of any form to the rapport—then it could be applied to material for which the ornamental language of the 19th century was never intended.Ornament is a challenging material for diffusion models. Base Stable Diffusion handles the depiction of objects and scenes well, but it struggles to maintain regularity: the rapport breaks down, the counting falters, and the elements lose uniformity. Training here imparts not a stylistic nuance, but a new skill, and the result can be objectively verified—one simply has to see whether the pattern coheres.

Source images for training
The training set consists of forty-four fragments from ornamental tables taken from «Ornament Grammar, ” sourced from Vikisklad. All images are in the public domain. The images were selected manually, not through automated sheet cropping. This method proved to be both more precise and faster: on Vikisklad, next to the full scans of the spreads, individual cut-out fragments are available—already framed according to the repeat pattern, without book margins or captions. Each image has been resized to a 1024 by 1024 pixel square. The selection is diverse in its geometry—ranging from dense geometric lattices to floral scrolls and weaves: in all cases, there is flat color filling, no shading, and a limited palette characteristic of chromolithography.
Description of the learning process
1. Dataset collection and verification of image legal status.
2. Resizing images to a square format and uniform resolution.
3. Manual annotation of signatures with trigger token input.
4. Environment setup and hyperparameter tuning.
5. Training the LoRA adapter while saving intermediate checkpoints.
6. Series generation and results analysis.
First, I connected Google Drive and verified the sample composition: the number of files and the uniformity of sizes.
The next cell compiles all dataset images into a single numbered grid. It serves two functions: it allows for an assessment of the overall sample homogeneity and provides a numbering system that I subsequently used to annotate each frame.


At this stage, I was annotating the training dataset. In the initial example, the caption is generated by a BLIP model trained to describe photographs—and on ornamental tables, it consistently outputs the same phrase for every frame: «a close up of a pattern.» This annotation does not distinguish between the images and therefore provides no training signal, so I discarded it and manually annotated the selection. The annotation principle is as follows: the caption describes what differs between the frames—the layout type, the nature of the motif, the dominant color. What is constant across the entire dataset is not mentioned and is tied to the trigger token ORNGRM, which is absent from the base model’s vocabulary. If constant features were described in plain language, the style would dissolve into them and could not be evoked separately.
Next, I installed the libraries required for LoRA training. Unlike the original example, where diffusers is installed from a development branch, I used the release version from PyPI, and I downloaded the training script from the tag corresponding to the installed version. The contents of the development branch change daily, making this run non-reproducible.
Here, I ran training using the DreamBooth method with a LoRA adapter.The training consisted of eight hundred steps at a 768 resolution. I increased the resolution from 512, which was set in the example, because SDXL was trained on 1024, and it loses detail at half resolution—detail that is critical for fine ornamental structures. Intermediate slices are saved every two hundred steps.Only the UNet was trained—fine-tuning the text encoder does not fit into memory and is not necessary for a style task.
Learning Dynamics
Intermediate weights allowed us to trace how the style developed. The images were generated using a single prompt and a single initial noise seed—only the number of steps taken varied.
Learning Model Dynamics
Prompt: «ORNGRM ornamental plate, interlaced star lattice»
The progression turned out not to be a gradual improvement, but a change in states. The characteristics were assimilated in a specific order: first, the typographic style, then the regularity, then the scale and color palette. Between the four-hundredth and six-hundredth step, the model restructured the grid’s very framework rather than refining the previous one, which suggests that by the eight-hundredth step, the learning had not yet plateaued.


Before and After Learning
Prompt: «ORNGRM hexagonal tile grid with rosettes»
Before training, when prompted for a hexagonal grid with sockets, the model produced a relief honeycomb surface with no sockets—part of the prompt was being lost. After training, the sockets appeared and were clearly defined as hexagons in each cell: the model learned to retain both parts of the prompt simultaneously, not just the dominant element.
Resulting Image Series
The first three narratives reproduce the ornament types presented in the training set. They serve to verify that the style has been absorbed. Each narrative is generated using three different seeds—numbers that define the initial noise from which the model develops the image. With a fixed seed, the differences are attributed to the wording, not to randomness, which allows us to separate the influence of the prompt from the influence of the starting point.
Plot 1
Prompt: «ORNGRM ornamental plate, interlaced star lattice»
The lattice structure and the star node are reproduced in all three variations, differing in scale and fill density. The first option yields a medium-scale rhombic grid with light gaps; the third is twice as dense and features diagonal hatching within the cells. The second option is more radical: a pair of dark purple and lemon tones is replaced by brick red on lilac, the paper texture disappears, the line becomes smooth, and the solid fill is closer to screen printing than to a lithographic print.
