NSTPAPER — A world crafted from paper
NSTPAPER — an experimental project focused on teaching Stable Diffusion XL to an authorship visual style based on multi-layered paper dioramas. The project is founded on the idea of translating the surrounding world into a conditional material reality, where people, architecture, plants, natural landscapes, and objects appear as if they were manually cut from colored paper and assembled from numerous separate layers. The main goal was not to teach the model to reproduce a specific object, but rather to convey a comprehensive visual language that could be applied to completely different subjects. Characteristic elements of NSTPAPER included the matte paper texture, the visible edges of cut elements, the physical thickness of the layers, simplified forms, and soft shadows that create the feeling of a genuine paper construction.
Visual Language:
Layering The image is constructed like a physical composition made from individual paper cutouts.
Materiality The surfaces retain the matte, uneven texture of paper and cardboard.
Paper Edges Object contours appear not drawn, but physically cut.
Soft Depth Shadows between layers create volume without striving for photorealism.
Form Simplification Complex real-world objects are transformed into cleaner, more graphic silhouettes.
Learning Material
For training the model, a dataset of 13 square images was created, all unified by a consistent visual style. When selecting the images, it was crucial to maintain variety in the subjects while ensuring the visual characteristics remained as consistent as possible. Thanks to this, the model should learn to associate the NSTPAPER trigger not with a specific depicted object, but with a system of visual features. The images in the dataset are based on my own photographs.
The Learning Process
Stable Diffusion XL 1.0 was used as the base model. To adapt the model to a unique style, the LoRA (Low-Rank Adaptation) method was employed, which allows for training visual nuances without fully retraining the original model.
For each image in the dataset, a text description was automatically generated using BLIP. A specific trigger, NSTPAPER, was added to the captions.
This became the designation for the new visual style of the model.
Training was conducted on 13 images at a resolution of 512 px over 500 steps.
Paper City
The architectural scene demonstrates how NSTPAPER interacts with a complex system of geometric objects. Facades, roofs, windows, and street elements are divided into individual planes. The style is particularly evident at the boundaries of the buildings: instead of realistic materials, paper surfaces appear, and space is formed through the layering and shadows between them. The architecture remains recognizable but takes on the quality of a model or set design.
The Person Within the Paper World
Character generation was necessary to test more complex organic forms. Unlike architecture, the human figure contains numerous smooth lines and minute details. The model preserves a legible human silhouette while simultaneously simplifying it and subordinating it to the overall graphic plasticity of the image. The environment and the character exist within the same material space—the person does not look like a photograph pasted over a stylized background.
Paper Landscape
The natural landscape proved especially suitable for the principle of layering. The mountains, the river, the trees, and the clouds naturally divide into planes. This allows the effect of a paper diorama to become part of the image’s spatial structure: the foreground, middle ground, and background are perceived as distinct physical layers.
When the model resists the style
Generating an individual still life became the most telling experiment. On the first attempt, the model tried to render the camera as a realistic object, showing plastic, metal, glass, and surfaces typical of still-life photography. Therefore, the prompt was further refined with descriptors like «colored paper, ” „cardboard, ” „visible paper edges, ” and „layered paper pieces, ” and realistic materials were added to the negative prompt. After this refinement, the still life began to align much better with the learned visual language.
What was the model able to learn?
The most consistent result was not a specific color palette or a particular type of object, but the underlying logic of the image construction itself. All works share the characteristic features of NSTPAPER: object layering, defined element edges, matte surfaces, simplified geometry, and soft contact shadows. However, the style manifests differently across categories. In architecture, geometric rigidity and the feeling of a model dominate. In nature scenes, the paper layers simultaneously function as spatial planes. In figurative imagery, the model must strike a balance between recognizable form and stylization. Object generation, conversely, demonstrated the boundaries of LoRA: the inherent knowledge of basic SDXL regarding the appearance of real objects sometimes competes with the learned style. Thus, the final series does not present a repetition of a single image, but rather the transposition of a unified visual principle across diverse categories of subjects.
Technical Details
The project includes a Google Colab Notebook that contains the complete process for dataset preparation, caption generation, Stable Diffusion XL configuration, LoRA training, and final model verification.[https://drive.google.com/file/d/1PTUTzWiq-nTefZXfqjdJJeecPGHzFAP7/view?usp=share_link]
Generative Model Application Description
The project utilized several artificial intelligence-based tools.
Stable Diffusion XL 1.0 served as the core generative model. Based on it, a custom model, NSTPAPER, was trained using the LoRA method to generate images in the developed style.
BLIP Image Captioning was used to automatically generate textual descriptions for the training dataset images.
ChatGPT (OpenAI) was employed as an auxiliary tool for code assistance, prompt formulation, and preparing the project’s textual description.
ChatGPT (OpenAI) was also used to transform original author photographs into stylized images, from which the training dataset was formed.













