Image Generation
SaaS, open source and skills for image generation work — with the limits of each written down.Before you pick
What decides it
Whether fixing one thing preserves everything already right decides this; otherwise each correction restarts the search for the frame you nearly had.
What to avoid
Commercial rights come from the licence and the plan tier, not from holding the image. Open weights and revenue thresholds quietly withdraw them.
What is changing
One-shot generation is levelling out; editing is where models now separate. Locking a face, product or style across a whole set is becoming the test.
SaaS
Hosted products you sign up for.Leading
Established picks, at the top of today's aggregate.
- 1GPT Image 2 (ChatGPT)Tops the Artificial Analysis and LMArena image arenas for prompt adherence and text. Best all-round pick.
- 2Nano Banana Pro (Google Gemini)Leads on photorealism, character consistency and UGC ad creatives. Tops CNET's 2026 ranking.
- 3MidjourneyStill the king of artistic style — concept art, moodboards, stylized campaigns (V8.1).
- 4FLUX.2 (Black Forest Labs)Strongest open-weight option. Go-to for product photography and embedding generation into an app.
- 5IdeogramBest at text inside images — logos, posters, typography. Generous free tier.
Emerging
Newer challengers, surfaced by two or more sources.
Open source
Repositories you run yourself.Leading
- 1Stable Diffusion WebUI (AUTOMATIC1111)The most-adopted local image-generation tool ever (~164k GitHub★) — the de-facto Stable Diffusion interface with a vast extension ecosystem.
- 2ComfyUIThe node-based industry standard (~122k GitHub★) for pro Stable Diffusion / FLUX pipelines; 12k+ custom nodes and full pipeline control.
- 3Fooocus (lllyasviel)The simplest local generator (~51k GitHub★) — a Midjourney-like, prompt-only experience on SDXL. Best on-ramp for non-technical users.
- 4Diffusers (Hugging Face)The standard library (~34k GitHub★) for running/embedding diffusion models (SD, SDXL, FLUX) into your own app — the developer default.
- 5InvokeAIClean, professional local studio (~27.7k GitHub★) with a powerful unified canvas — combines SD WebUI and ComfyUI strengths for creative workflows.
Emerging
- 1jd-opensource/JoyAI-ImageOpen-weight model family from JD.com pairing an 8B understanding model with a 16B diffusion transformer for generation and instruction-based editing.
- 2open-mmlab/mmagicPyTorch toolbox from OpenMMLab bundling 50+ research models for text-to-image, inpainting, matting, super-resolution and video restoration in one API.
- 3mylxsw/aideaOpen-source Flutter app for phones and desktop that fronts many chat and image models, with its own self-hostable backend server.
- 4kadevin/ilab-conjureLocal-first workbench for GPT-Image and Gemini APIs, adding a shared reference gallery, prompt templates, reusable chips and a concurrent job queue.
- 5kingbootoshi/nano-banana-2-skillCommand-line tool and Claude Code plugin for Gemini image models, with green-screen transparency, style transfer, batch runs and per-image cost tracking.
Skills
Skills you drop into a coding agent.Leading
- 1nano-banana-pro (intellectronica/agent-skills)The most-cited image Agent Skill — generate/edit images up to 4K with Nano Banana Pro (Gemini 3 Pro Image), strong on legible in-image text.
- 2ccskill-nanobanana (feedtailor)Claude Code image-generation skill on the Gemini 3 Pro Image (Nano Banana Pro) model. Clean, well-documented but single-author.
Emerging
- 1wuyoscar/GPT-Image2-SkillA CLI and agent skill for OpenAI's gpt-image-2 generation and edit endpoints, shipped with a categorised gallery of worked example prompts.
- 2guinacio/claude-image-genClaude Code skill plus MCP server routing prompts to Gemini or gpt-image-2, saving files locally with up to five reference images.
- 3shinpr/mcp-imageMCP server for Gemini, GPT Image and Seedream that rewrites short requests into detailed photographic prompts and offers speed-versus-quality presets.
- 4hassancs91/claude-image-generation
- 5wanshuiyin/ARIS-Movie-DirectorAgent pipeline that turns a rough story into a storyboarded image sequence, with separate models blind-reviewing each frame before it ships.