Ideogram 4.0 is the current Ideogram model and the first in the line to ship as open weights. What changes is narrow: you describe a layout, including the words that have to come out readable.

This guide covers where 4.0 runs, how a first image comes together, one worked prompt, and what Ideogram has not published.

What Ideogram 4.0 is

The official model page defines it in a sentence. Ideogram 4.0 is the current Ideogram model for prompt fidelity, crystal-clear type, reliable editing, native transparency, style control and design-quality outputs.

The line under it explains the feature. The model was trained with a describe-to-structure-to-recreate loop, reading scenes, text and objects as structured data before rebuilding images from that representation. Bounding boxes were taught alongside plain-language descriptions, so a headline or a product can be told where to sit.

Documented capabilityWhat the official pages say
Text in the frameIdeogram has led on text rendering since launch, and 4.0 adds bounding-box layout control
Layout controlPositions come from bounding boxes trained with plain-language descriptions
Editable outputTransparent cutouts and editable text layers already ship
Open weightsPublished to download, fine-tune and run on your own hardware
Earlier models3.0, 3.0 (March 26), 2a, 2.0, 1.1 and 1.0 stay selectable

Ideogram says generations arrive as editable files rather than flat frames.

Where you can run it today

The docs list four approved public surfaces for 4.0: the Ideogram app, API workflows, MCP for agents and internal tools, and open weights.

SurfaceWhat the official pages attach to it
Ideogram appCreate in the browser or the mobile app, on a subscription
APIEndpoint schemas live at developer.ideogram.ai, billed per image
MCPListed for agents and internal tools
Open weightsDownload and self-host, under a licence that scales with deployment

One split catches people out. The docs state that subscriptions and API accounts are separate, with separate payment and billing, so a plan bought in the app does not fund API calls.

The Ideogram API guide covers the request shape and that billing split.

Your first image, step by step

Official Ideogram 4.0 example: a CONCHA tasting menu card on dark green cloth, its small print legible
Ideogram

The first run confirms the loop; it is not the run for your best image.

  1. Create the account. Email signup works, and the quick-start guide warns it includes no free credits.
  2. Treat the username as permanent, like the sign-in method and email address.
  3. Keep the prompt under about 150 words, around 200 tokens, and lead with the subject.
  4. Set model, ratio, style and references, and check the model you picked.
  5. Generate, review, download. Types include JPEG, PNG or WebP, and transparent output disables JPEG.

Aspect ratios are presets, from tall 1:3 and 9:16 through 1:1 to wide 3:1, and API requests write 16:9 as 16x9. Style Reference takes up to three images in Auto, General, Realistic or Design mode.

Plain text is supported. Magic Prompt can expand a natural language prompt into a full JSON caption automatically, which is how a sentence typed in the app reaches a model trained on structured captions.

Worked example: a product shot with two lines of type

The job is a coffee bag on a counter with a legible product name and one short line under it. Complex frames work against legibility, so the scene stays plain and the words few.

A matte black coffee bag standing on a concrete counter, studio light from the left.
Text on the front of the bag, in a clean sans-serif, reading "MORNING BLEND".
A smaller line underneath reading "Small batch".
Warm grey background, soft shadow, product photograph, shallow depth of field, 4:5.

Each line answers a different question. The words arrive early, quoted and short, which is three of the five practices the official typography guide lists; the other two are visual description and low complexity.

The result splits along the same limits. The front line is two words, so it survives. The second line is where length costs accuracy, because longer text is more likely to be misspelled.

Case two moves one block. “Same two lines in the same place. Replace the studio with daylight through a shop window and the counter with a wooden table.” The type is untouched, so the change tests whether a layout holds when the scene moves.

Order matters more than length, and that goes for the structured path too: the official JSON guide states that key order matters.

A prompt structure that holds up

Both cases share a skeleton, which is what makes a good result repeatable rather than lucky.

The parts worth writing every time

Ideogram draws picture and lettering in one pass, so a prompt that never names the words hands the headline to the model. The blocks below come from the official prompting guide.

Subject: what is in frame, described in visual terms
Text: the exact words, quoted, near the front of the prompt
Placement: where the words and the main object sit
Style: photo or art style, with medium, lighting and palette if you have them
Format: aspect ratio, and the resolution tier
Keep out: what should not appear, written as a description of what should
Constraints: what must stay identical when you reuse a reference

Where text rendering breaks

The guide names three limits and one surprise. Longer text is more likely to be misspelled. Document-level typesetting is not what the model is for. Text rendering is most accurate in English; non-Latin scripts are unpredictable. The surprise: it is not possible to specify a typeface by name, so you describe lettering rather than request a font.

Misspellings have documented exits. Regenerate; shorten the word; overtype it in the Canvas text tool, which the docs present as the answer when a prompt cannot replace text; or rebuild the layout from a reference image.

Limits, cost and what Ideogram has not published

Upscale is paid, doubles the resolution at most, and crops to the nearest supported ratio. Output resolution is messier: the model page says realistic 2K images but publishes no pixel table, and the docs say exact pixels depend on model, size tier and endpoint.

Cost is where this article keeps its hands off the numbers. The official pages price the API per image across three quality tiers and keep plan and credit amounts on the live pricing page, and the docs warn against relying on static tables for purchase decisions. The shape is stable: API calls bill per image, and subscriptions and API accounts are separate.

One ranking note: the model page cites third-party benchmark results, DesignArena among them, which is a vendor-transcribed third-party result rather than an independent audit.

The Ideogram topic page collects the rest of our coverage, and the GPT Image 2.5 vs Ideogram comparison takes both models apart on typography and deployment.

Frequently asked questions

Do I have to pay to try Ideogram? The docs publish no free-credit amount and point at the live pricing page instead. Email signups, they do say, include no free credits.

Can I keep what I make? Downloads arrive as JPEG, PNG or WebP, and transparent output disables JPEG. Images made through the API have expiring links, so save them yourself.

Why does my text come out misspelled? Longer strings fail more often, and non-Latin scripts are unpredictable. Shorten the line, keep it in English, quote it.

Can I write the prompt in Chinese? Natural language is supported, and Magic Prompt expands it into a full JSON caption. Whether the lettering lands is another question, and the guide is blunt: English is most accurate, non-Latin scripts are not. Our lettering guide covers the quoting rules.

Get the order right once, words first and quoted, and the model stops being a slot machine.