Original size 2480x3500

Training a Generative Model Based on Ancient Chinese Script

PROTECT STATUS: not protected
Longread translated automatically
The project is taking part in the competition

Concept

The Oracle Project aims to create a generative model that transforms the aesthetics of ancient Chinese oracle bones (Jiaguwen) into modern minimalist icons and graphic design elements.

Generation creates new opportunities for developing stylized modern imagery, inspired by ancient art, in the fields of UI/UX design, branding, and identity.

Image Series

To create the dataset, I compiled a selection of 198 square images (1:1, 512×512 after preprocessing: center crop + contrast) depicting elements of oracle bone script (ancient Chinese writing).

big
Original size 3500x971

examples of source images

The Learning Process

The learning process began with checking for the presence of a GPU (Tesla T4) and installing the necessary libraries (diffusers, transformers, accelerate, peft, bitsandbytes).

The dataset was not compiled through a direct upload to Colab but via Google Drive: the images were pre-uploaded to a separate folder on Drive, and the notebook mounted Google Drive to copy them into the working directory. In the end, the dataset consisted of 198 images in the style of ancient Chinese oracle bone script, all standardized to a 512×512 format (center crop to square + contrast enhancement).

The BLIP model (Salesforce/blip-image-captioning-base) was used to generate text descriptions for each image. A fixed style identifier prefix, «ancient Chinese oracle bone script symbol, ” was prepended to each automatically generated description—this token is what links the content of a specific image to the style overall. All „file—description“ pairs were saved in metadata.jsonl.

Training was conducted using the DreamBooth + LoRA method based on Stable Diffusion 1.5 (stable-diffusion-v1-5/stable-diffusion-v1-5) with Hugging Face’s official script train_dreambooth_lora.py. Key parameters included a resolution of 512×512, train_batch_size=1 with gradient_accumulation_steps=2, a learning_rate of 1e-4, 480 training steps with checkpoints saved every 120 steps. The training took approximately 21 minutes on the T4.

After training, the LoRA weights (file pytorch_lora_weights.safetensors, ~3.1 MB) were saved locally, archived, and downloaded to the computer; an optional feature allows uploading them to the Hugging Face Hub repository.

Additional experiments were also performed: a comparison between the baseline (SD 1.5 without LoRA) and the fine-tuned model using a single prompt—to visually demonstrate the contribution of the fine-tuning; as well as a two-stage generation process (base object image without LoRA $\to$ style transfer via img2img) as a technique for subjects where direct generation with LoRA resulted in the loss of either object shape or stylistic texture.

Final Image Series

0

Using Genii

— Deepseek was used to refine the code by requiring fewer training iterations.

Training a Generative Model Based on Ancient Chinese Script
Project created at 26.09.2026