Image posted by CivitaiOfficial

default creator card background decoration

CivitaiOfficial

Uploaded

Process

Tools

Techniques

txt2img

Generation data

COPY ALL

Resources used

OpenAI's GPT-image-1
Checkpoint
4o Image Gen 1

Prompt

External Generator

A wide image taken with a phone of a glass whiteboard, in a room overlooking the Bay Bridge. The field of view shows a woman writing, sporting a tshirt wiith a large OpenAI logo. The handwriting looks natural and a bit messy, and we see the photographer's reflection. The text reads: (left) "Transfer between Modalities: Suppose we directly model p(text, pixels, sound) [equation] with one big autoregressive transformer. Pros: * image generation augmented with vast world knowledge * next-level text rendering * native in-context learning * unified post-training stack Cons: * varying bit-rate across modalities * compute not adaptive" (Right) "Fixes: * model compressed representations * compose autoregressive prior with a powerful decoder" On the bottom right of the board, she draws a diagram: "tokens -> [transformer] -> [diffusion] -> pixels"

Discussion

Civitai memberships — more features, fewer limits.