Browser

ChatGPT Image Handoff

Try it

Route a multi-image content job through the best available generation backend, prepare a resumable prompt pack with readable-text fact boundaries, optionally...

What it does

Route a multi-image content job through the best available generation backend, prepare a resumable prompt pack with readable-text fact boundaries, optionally...

The skill document

ChatGPT Image Handoff

Turn an image-generation gap into a recoverable local workflow. Default to backend=auto, preserve one job per image even when several jobs share one ZIP delivery, and converge every backend on the same import, QA, and human-confirmation contract.

Inputs

Collect:

  • a JSON spec with run_id, shared visual direction, dimensions, and ordered image jobs
  • one self-contained prompt and stable intended filename per job
  • approved reference-image paths when needed
  • a local handoff directory, preferably under an ignored runtime folder
  • backend, defaulting to auto
  • preferred_backend, defaulting to chatgpt_computer_use
  • per-job fact_sensitivity: low, medium, or high
  • per-job render_strategy: auto, generative, mixed, or deterministic
  • optional text_policy with mode, allowed_text, forbidden_patterns, and extra_readable_text
  • an existing jobs.json when resuming

Require Python 3 and Pillow for the bundled import and contact-sheet scripts. Route incomplete copy, unstable claims, or missing source anchors back to the upstream planning skill.

Backend Decision

Treat backend selection as runtime capability routing, not as an installation assumption.

BackendSelect whenHandoff
chatgpt_computer_useComputer Use exists and the active ChatGPT session is authenticatedUse ordinary Chat one job at a time; test ZIP only when native archive delivery is already supported
built_in_imagegenThe built-in generator is available and healthyReturn the queue to resilient-imagegen for serial generation
manual_chatgpt_handoffComputer Use is unavailable or unsuitable and ChatGPT use is allowedGive the prepared pack to the user, then resume at download import
mixedGenerated visuals are useful but exact text or structure needs deterministic overlayGenerate backgrounds, then finish in cards-to-images
deterministicExternal ChatGPT transmission is declined or exact text and diagrams dominateRender locally in cards-to-images

For backend=auto, try the preferred backend first when viable, then use this default order:

  1. chatgpt_computer_use
  2. built_in_imagegen
  3. manual_chatgpt_handoff
  4. mixed or deterministic

Do not silently choose a paid CLI or API path. Leave that as a separately confirmed resilient-imagegen fallback.

Choose the render strategy per job before selecting the runtime backend:

  • keep low-density conceptual visuals generative
  • use a strict allowlist for readable names, numbers, filenames, labels, or claims
  • recommend mixed for fact_sensitivity=high when a local renderer is available
  • let generation create the background and composition in mixed; add approved copy deterministically
  • record an explicit generative override when the user prefers full image generation, then require strict allowlist QA

Keep visual composition high freedom. Keep readable factual text low freedom.

Workflow

1. Prepare The Pack

Run:

python /scripts/prepare_handoff.py \
  --spec  \
  --out 

Confirm the pack contains:

  • master-prompt.md
  • jobs.json
  • prompts/.md
  • inbox/
  • imported/
  • reports/
  • reports/text-overlay-plan.json for mixed or deterministic jobs

Treat jobs.json as the source of truth. Keep this runtime directory out of public commits because it may contain local paths and reviewed working files.

For a resumed run whose pending or failed prompts must inherit a revised fact policy, run:

python /scripts/prepare_handoff.py \
  --spec  \
  --out  \
  --resume \
  --refresh-prompts

Refresh only queued, failed, or needs_revision jobs. Preserve completed images and invalidate external approval for every changed prompt.

2. Inspect Capabilities

Determine actual runtime state:

  • mark Computer Use available only when its tool and skill are exposed
  • read the active browser state to classify ChatGPT as authenticated, login_required, or unknown; do not assume a session from the machine or project
  • mark built-in ImageGen healthy only when the generator exists and is not in a known failed state
  • mark the local renderer available when cards-to-images can produce deterministic or mixed output
  • record whether ChatGPT transmission is pending, approved, or denied

Resolve and persist the backend:

python /scripts/resolve_backend.py \
  --handoff  \
  --computer-use available \
  --chatgpt-session authenticated \
  --built-in-imagegen unhealthy \
  --local-renderer available \
  --chatgpt-transmission pending

Use unknown when a capability has not been checked. Do not turn uncertainty into a positive capability claim.

3. Execute The Selected Branch

For built_in_imagegen, hand the queue to resilient-imagegen and import each selected local output before starting the next job.

For deterministic or mixed, hand the approved card or illustration package to cards-to-images; keep the same stable filenames and job ids.

When a generated image already has a useful composition but contains unapproved readable text, do not keep regenerating the full card. Create a reviewed mixed-overlay plan whose text elements exactly match the job allowlist, then render it locally:

python /scripts/render_mixed_overlay.py \
  --source  \
  --plan  \
  --output  \
  --report 

The renderer locks the reviewed source checksum, can normalize a matching-ratio source when the plan explicitly sets source_fit=resize_matching_ratio, makes source text unreadable with a declared blur-and-tint treatment, rejects text outside allowed_text, and can require every approved string to appear exactly once. Import the derived image with --generation-backend mixed --replace, then repeat strict visual QA. Keep layout choices in the reviewed plan; do not use this renderer as automatic OCR or an unreviewed redaction tool.

A mixed result must retain a materially visible and relevant generated visual contribution. Review the full generated card first. When incorrect readable text is confined to one or two clear regions and the rest of the composition is strong, preserve the accepted pixels and patch only those regions with reviewed copy. Use a text-free background with reserved copy zones when text failures are widespread or cannot be isolated safely. Do not apply blur or tint strong enough to erase the composition or series style. If the generated source becomes visually incidental after the overlay, classify and report the result as deterministic, not mixed. Treat visual continuity with already accepted series images as part of strict QA.

A full-card generative attempt may be finished as mixed after human review when its failures are localized. Keep the requested render strategy and the final per-image generation backend as separate provenance fields so the repair does not rewrite history.

For a strict generative job, send the approved readable-text allowlist with the prompt. Permit creative visual metaphors while rejecting every additional readable label, identifier, metric, hash, checksum, timestamp, version, or path.

For manual_chatgpt_handoff:

  1. Give the user master-prompt.md once.
  2. Give one prompts/.md at a time.
  3. Ask the user to place downloaded results in the chosen downloads folder or inbox/.
  4. Resume from detection and import without rebuilding completed jobs.

For chatgpt_computer_use:

  1. Read and follow the installed computer-use skill.
  2. Use its persistent JavaScript runtime for every GUI action.
  3. Fetch fresh app state before every decision and after every action.
  4. Prefer accessibility elements; use screenshot-guided coordinates only when required.
  5. Stay in the user-approved authenticated browser window and ordinary Chat surface. Do not require a specific Chrome profile, project, or Work mode.
  6. Use one job at a time by default. Consider one ZIP batch only when the active surface has demonstrated native multi-image archive delivery; preserve separate job ids and files.
  7. Never reuse stale element indices.
  8. Verify a new local image or ZIP archive before marking any job downloaded.

For an optional ChatGPT ZIP batch, read ZIP batch delivery, then prepare the reviewed batch:

python /scripts/batch_zip.py prepare \
  --handoff  \
  --batch-id  \
  --job-id  \
  --job-id 

Send only the generated batches//work-prompt.md. Require separate exact filenames, qa-report.json, batch-manifest.json, and one ZIP. Treat ZIP delivery as an experimental optimization of browser interaction, not a relaxation of per-image QA. A misclassification as missing-image editing, a programmatic-rendering proposal, or a missing archive is a batch failure; record it with batch_zip.py fail and return to ordinary Chat one job at a time.

When the user requires ChatGPT-native image generation, reject Python, Pillow, SVG, HTML, Canvas, deterministic drawing, screenshots, or placeholders inside ChatGPT as backend substitution. Record the batch as failed and retry fewer jobs through the native image tool.

Before the first ChatGPT submission, list the exact prompt and reference files, destination, and purpose. Request action-time confirmation immediately before typing the first prompt or uploading any reference. One confirmation may cover the unchanged reviewed batch. Stop for login, account verification, CAPTCHA, quota, billing, or an unexpected destination.

Do not encode selectors, coordinates, conversation ids, account identifiers, or active-session details in public artifacts.

4. Detect And Import Outputs

Scan recent image downloads without changing them:

python /scripts/detect_downloads.py \
  --handoff  \
  --downloads 

Import one verified result:

python /scripts/ingest_images.py import \
  --handoff  \
  --job-id  \
  --source 

Copy the original into inbox/, create the stable output under imported/, record checksums and dimensions, and refresh the contact sheet. Refuse a different replacement unless --replace is explicit. Never delete the original download. Pass --generation-backend mixed or --generation-backend deterministic when a local derived output overrides the run-level backend.

Import one reviewed ZIP archive:

python /scripts/batch_zip.py import \
  --handoff  \
  --batch-id  \
  --archive 

Preserve the ZIP, reject unsafe entries, import only manifest filenames, and leave missing files resumable. Never trust the archive's qa-report.json as a local QA pass.

5. Inspect And Record QA

Open every imported image. Check:

  • subject, composition, order, dimensions, and intended ratio
  • exact Chinese, numbers, filenames, arrows, and Skill names
  • mobile-size legibility, safe margins, cropping, and consistent series direction
  • invented claims, logos, private paths, watermarks, or fake interfaces
  • readable text outside text_policy.allowed_text
  • any match or semantic equivalent from text_policy.forbidden_patterns
  • whether the result should become a visual background with deterministic text overlay

Record QA:

python /scripts/ingest_images.py qa \
  --handoff  \
  --job-id  \
  --status passed \
  --text-policy-status passed \
  --notes "checked at mobile size"

Use needs_revision when the image fails. Route exact typography and source-backed diagrams to deterministic or mixed rendering instead of repeatedly trusting generated text.

Treat an invented checksum, hash, commit id, timestamp, version, metric, filename, or path as a factual failure even when it looks decorative. For strict jobs, compare all readable text with the stored allowlist before passing QA.

For a strict failure, add each extra phrase with --observed-extra-text and each matched deny item with --matched-forbidden-pattern. The script must refuse status=passed until text_policy_qa=passed with no recorded extras.

6. Resume And Hand Off

Read jobs.json after any interruption. Resume only queued, failed, or needs_revision jobs. Read UI recovery rules before retrying a browser, session, or download failure.

For ZIP delivery, also read zip_batches. and batches//batch-import-report.json. Resume only the listed missing jobs; do not regenerate images that were imported successfully.

When all images pass QA:

  • hand imported/ and reports/contact-sheet.png to cards-to-images or article-to-illustrations
  • update the downstream image manifest with generation_backend and the handoff run_id
  • keep human_confirmation=pending
  • route to md-img-r2 plan mode only when stable public URLs are needed

Record an approval only after the user has reviewed the images:

python /scripts/ingest_images.py confirm \
  --handoff  \
  --job-id  \
  --status approved

State Contract

Use job states queued, prompt_submitted, generating, downloaded, imported, qa_passed, needs_revision, and failed.

Store these policy fields on every prepared job:

fact_sensitivity: high
requested_render_strategy: auto
render_strategy: mixed
recommended_render_strategy: mixed
render_strategy_reason: high fact sensitivity with an available deterministic text renderer
text_policy:
  mode: allowlist
  allowed_text: []
  forbidden_patterns: []
  extra_readable_text: reject
text_policy_qa:
  status: pending
  observed_extra_text: []
  matched_forbidden_patterns: []

Keep this runtime block in jobs.json:

generation_runtime:
  requested_backend: auto
  preferred_backend: chatgpt_computer_use
  computer_use: unknown
  chatgpt_session: unknown
  built_in_imagegen: unknown
  local_renderer: unknown
  chatgpt_transmission: pending
  selected_backend:
  selection_reason:
external_handoff:
  destination: ChatGPT
  status: not_started
zip_batches: {}

Require a local file, checksum, dimensions, and visual QA before claiming success. Do not infer completion from generator UI text.

Boundaries

  • Keep account and session state local; publish only the generic workflow and scripts.
  • Do not automate account creation, billing, CAPTCHA handling, R2 upload, URL rewriting, or platform publishing.
  • Do not send sensitive or unreviewed material to ChatGPT.
  • Do not overwrite reviewed images silently.
  • Stop at human image confirmation before any external publication.

作者入口

Related skills

Detect image generation requests and route them to the artist agent. Use when user asks to draw, generate, create, sketch, render, or visualize an image.

Stabilize multi-image generation by converting prompts into a retryable serial job queue, inspecting runtime capabilities, routing through built-in ImageGen,...

1 installs

Generate raster images (PNG/JPEG/WebP) using the user's ChatGPT subscription via a local one-file Python CLI — no OPENAI_API_KEY, no gateway, no daemon. Two...

1 installs

Turn image briefs into production-ready GPT Image 2.5 Studio jobs with exact-text manifests, reference-image role contracts, ratio and resolution choices, controlled edit rounds, and visual QA. Use for posters, packaging, ecommerce images, UI mockups, storyboards, localized creatives, or precise image edits in the GPTImage-2-5.com web Studio; do not use for generic prompt lists or undocumented API integration.

Generate and edit images with GPT Image through RunAPI. Use when the user asks an agent to create, edit, or transform images with GPT Image. Default to the RunAPI CLI for one-off generation; use SDKs only when the user is integrating RunAPI into an app or backend.

13 installs

Multi-modal content creation workflow for freelance creators. Receives customer requests via WhatsApp (text or voice note), transcribes audio with the OpenAI...

3 installs