Krea 2: An open-weights 12B image model tuned for creative range, not safe defaults
Krea has released Krea 2, a 12-billion-parameter open-weights image generation model built around a contrarian premise: most production diffusion models have converged on a narrow band of polished, default aesthetics, which makes them reliable but poor at creative exploration. Krea 2 instead optimizes for a broad visual space and gives users steering controls to move through it. The model is a diffusion transformer trained through a multi-stage pipeline — pretraining, midtraining, supervised finetuning, preference optimization, and reinforcement learning — and incorporates architectural choices like grouped-query attention, sigmoid-gated attention, and integration of Qwen3-VL and iREPA to speed convergence and stabilize training. Krea reports a top-10 placement on the Artificial Analysis text-to-image leaderboard and second among independent labs.
The data philosophy is the most distinctive part. Rather than filtering aggressively for ‘high quality’ via aesthetic and image-quality-assessment scorers — which Krea argues bakes in bias by, for example, discarding intentional motion blur — the team keeps a diverse, broadly representative mix and removes only duplicates, mis-captioned samples, harmful biases, and overly complex low-resolution cases. Notably, they exclude all AI-generated images from pretraining, on the finding that even small amounts of synthetic data are easier for the model to learn and effectively cap output quality. Captioning runs through OCR plus a vision-language model enriched with metadata, then gets reformatted into varied prompt lengths, with training weighted toward long, dense captions for faster convergence.
To close the gap between richly-captioned training data and the short, vague prompts real users write, Krea ships two control systems: a prompt expander that enriches underspecified prompts without overriding intent, and a style-reference system that lets users transfer the mood or look of reference images with controllable strength and minimal content leakage. Pretraining follows a resolution curriculum (256px to 1024px), spending most compute at low resolution to build core capabilities before adding high-fidelity detail. Together these position Krea 2 as a foundation model aimed at exploration rather than one-shot polished output.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.