WayToClawEarn
Medium impactBlack Forest Labs Docs

Black Forest Labs Flux 3 Image: Single-Endpoint Multi-Step Local Edits, Bounding Box Layout, Up to 4K

Black Forest Labs released Flux 3 Image on October 1, 2026 — a unified image generation and editing model with a single endpoint where the prompt itself determines whether to generate or edit. Features include bounding box layout composition via JSON in the prompt, pixel-exact local edits that preserve the rest of the image, up to 10 reference images, 15 aspect ratios, and up to 4K output. API pricing ranges from $0.041 per image (768sq) to $0.607 (4K).

Edisen Lu · WayToClawEarnVia Black Forest Labs DocsPublished Oct 2, 2026

Reviewed from public sources · AI-assisted drafting under editorial oversight. How we work · Original source

Public-source compilation

Synthesized from public posts/docs. Prefer the original source for primary claims.

How we review content · Primary source · Black Forest Labs Docs

TL;DR

Black Forest Labs released Flux 3 Image on October 1, 2026 — the image pillar of the Flux 3 multimodal model family. The core innovation is a "single endpoint" design: the same API endpoint handles both image generation and image editing, with the prompt content automatically determining the workflow — no mode parameter needed. Features include bounding box layout composition, pixel-exact local edits (change one element, preserve everything else), up to 10 reference images, 15 aspect ratios, and up to 4K output (~16 megapixels). API pricing ranges from $0.041 per image (768sq) to $0.607 (4K).

Single-Endpoint Design: Why This Matters

Traditional image APIs split generation and editing into different endpoints or mode parameters. Flux 3 Image's innovation: the prompt itself determines the operation type.

  • Want to generate a new image? Write a descriptive prompt
  • Want to edit an existing image? Specify the element to change in the prompt
  • Want bounding box layout? Append a JSON box list to the prompt

No mode field, no endpoint switching. This means developers can handle both "user uploaded an image to edit" and "user wants a new image" with the same API call — significantly simplifying application logic.

Core Features Deep Dive

1. Bounding Box Layout Composition

Add a JSON-formatted bounding box list to the end of your prompt to precisely position each element:

  • Place a headline in the top-left
  • Locate a specific face in a crowd
  • Create a grid panel layout

Bounding boxes are written directly in the prompt text — no separate parameter field. This design makes layout information part of the prompt's semantics, helping the model better understand element relationships.

2. Pixel-Exact Local Edits

Mark the element to change and keep the rest of the image untouched:

  • Recolor: Change colors
  • Replace: Swap content
  • Move: Adjust position
  • Resize: Change scale
  • Remove: Delete elements

Key: Multiple elements can be edited in a single request. This isn't simple inpainting — the model understands semantic relationships between elements and automatically adjusts lighting, shadows, reflections, and other correlated aspects when editing one element.

3. Multi-Reference Image Composition

Supports up to 10 reference images:

  • Reference each image by its position in the prompt (e.g., "use image 1's style, image 3's composition")
  • Restyle an existing image
  • Combine elements from multiple references into a new composition

4. Resolution and Aspect Ratios

ResolutionApproximateUse Case
768sq~590K pixelsFast previews
1k~1M pixelsStandard quality
2k~4M pixelsHigh quality
4k~16M pixelsProfessional grade

Supports 15 aspect ratios covering everything from square to widescreen to vertical.

API Pricing

ResolutionPriceNotes
768sq$0.041Most economical, good for batch previews
1k$0.048Standard quality
2k$0.100High quality
4k$0.607Top resolution

Per-image pricing; resolution is the only price variable. No token-based billing, no extra parameter fees.

Competitive Landscape

At launch, Ideogram also released version 4.5 with a similar editing focus and plans for open weights. Key differences:

DimensionFlux 3 ImageIdeogram 4.5
Core selling pointSingle-endpoint unified generation + editingPrecise editing without damaging original
Bounding boxNative supportNot mentioned
Max referencesUp to 10Not specified
Max resolution4K (~16MP)Not specified
Open weightsComing in weeksComing in future
API pricing$0.041-$0.607/imageNot specified

Black Forest Labs' Flux series has been a benchmark in open-source image models. An open-weight version of Flux 3 Image is expected in the coming weeks.

Developer Integration Guide

API Call

All operations go through a single endpoint:

POST /api-reference/utility/generate-an-image-with-flux-3

The request body only needs prompt and resolution parameters. Whether to generate or edit, whether to use bounding boxes — all determined by the prompt content.

Typical Workflows

Scenario 1: Generate new image

code
prompt: "A serene mountain landscape at golden hour, snow-capped peaks"
resolution: "2k"

Scenario 2: Local edit

code
prompt: "Replace the red car with a blue bicycle [image: user_uploaded.jpg]"
resolution: "2k"

Scenario 3: Bounding box layout

code
prompt: "A magazine cover layout [
  {box: [0.1, 0.05, 0.9, 0.15], text: "headline"},
  {box: [0.05, 0.2, 0.45, 0.5], image: "portrait"},
  {box: [0.5, 0.2, 0.95, 0.95], image: "landscape"}
]"
resolution: "4k"

Use Case Analysis

Best-Fit Scenarios

  • E-commerce product image editing: Recolor, change backgrounds, remove unwanted elements while preserving the product
  • Marketing material batch generation: Same layout template, batch-replace text and image elements
  • Social media content creation: Quickly generate images in different aspect ratios
  • UI/UX prototyping: Use bounding boxes to quickly lay out interface elements
  • Multi-reference composition: Fuse styles and compositional elements from multiple product images into new visuals

Less Suitable Scenarios

  • Complex edits requiring fine-grained control over each reasoning step (consider multi-step API call chains)
  • Real-time interactive applications requiring ultra-low latency (image generation is inherently a seconds-level operation)
  • Fully offline use cases (API-only for now; open-weight version coming later)

Industry Context

Flux 3 Image is part of Black Forest Labs' broader multimodal strategy. The Flux 3 model family aims to unify image, video, and audio generation based on a "Self-Flow" joint learning architecture. Flux 3 Image is the image pillar of this vision, with video and audio modules expected to follow.

The image generation API market is competitive in 2026, but Flux 3 Image's "single endpoint + bounding box + multi-reference" combination offers unique advantages in developer experience — especially for applications that need to handle both generation and editing, where the single-endpoint design can significantly simplify code architecture.

Developer Action Items

  1. Try the free demo: Black Forest Labs offers an online editing demo to experience multi-step editing directly
  2. Evaluate API integration: If your app needs both generation and editing, the single-endpoint design can replace existing multi-endpoint logic
  3. Watch for open weights: If you need offline deployment or fine-tuning, wait for the open-weight release
  4. Compare pricing: $0.041/image (768sq) is in the economical tier for image APIs — suitable for high-volume generation scenarios
  5. Test bounding boxes: If your app involves layout composition (posters, magazine covers), prioritize testing the bounding box feature

Sources: Black Forest Labs Release Notes, Flux 3 Image API Reference, Bounding Box Tutorial, Online Editing Demo, The Decoder Coverage

Black Forest LabsFlux 3 Imageimage generationimage editingAI APIbounding boxmultimodal AItext-to-image

View source →

Educational reference only: cases summarize public sources and may use AI-assisted drafting under editorial review. Not financial advice; outcomes are not guaranteed.