Cookie Settings

    We use cookies to improve your experience on our website. You can choose which cookie categories you want to accept. Learn more

    Responsible Party
    Contact Form
    uNaice
    As of September 2026 · 85+ models · tokens & subscriptions compared

    Which AI model
    fits your task — and what does it really cost?

    Tell us what you want to do — we'll show you the right model, the honest token bill and why ChatGPT Plus, Claude Pro & co. are not directly comparable.

    This comparison updates automatically once a week · next update: Mon, 21/09/2026 · last updated: 14/09/2026
    Interactive model finder

    Find the right AI model in just a few clicks

    Pick your use case and three requirements — we suggest the best models and compare the costs for you.

    1

    Pick the use case that fits best.

    2

    Tells us whether to prefer cheap or top-tier models.

    Multi-select. For sensitive data usually EU (GDPR).

    The “context window” — the AI's short-term memory.

    Auto-matched to your selected use case – you can override anytime.

    Prices & context sizes automatically once a week updated · next update: Mon, 21/09/2026
    Your result

    Your personal recommendation

    Out of 40+ models, freshly calculated — this fits you best:

    Costs estimated for 500k input + 200k output tokens / month (≈ a few thousand requests). Need an exact calculation? Let's run the numbers together →

    DeepSeek

    deepseek-flash

    Top pick

    Du kannst es für schnelle, kostengünstige Aufgaben wie JSON-Ausgabe und Tool-Nutzung verwenden, wenn Geschwindigkeit entscheidend ist.

    Input: $0.006/1MOutput: $0.012/1MContext: 1M
    Cost / month (API)$0.01

    Google

    Gemini 2.5 Flash

    RAG, translations, image recognition at scale — excellent price/performance with long context.

    Input: $0.35/1MOutput: $1.05/1MContext: 1M
    Cost / month (API)$0.39

    ByteDance Doubao 🇨🇳

    Doubao Pro 1.5

    Extremely cheap for mass content, social tagging and content moderation.

    Input: $0.11/1MOutput: $0.28/1MContext: 256k
    Cost / month (API)$0.11

    Google

    Gemini 2.5 Flash-Lite

    Tagging product data, bulk translations, simple classification — when every cent per call matters.

    Input: $0.1/1MOutput: $0.4/1MContext: 1M
    Cost / month (API)$0.13

    Don't want to decide yourself?

    We build you multi-model routing that automatically picks the cheapest good-enough model per task — and integrate it into your ERP, CRM or PIM. Why “chat only” isn't enough →

    Free intro consultation

    Immer up-to-date bei neuen KI-Modellen oder Preisen

    Kurze Mail, sobald ein neues Modell erscheint oder sich Preise ändern. Kein Spam, jederzeit abbestellbar.

    Mit dem Absenden bestätigst du, dass wir dir Updates per E-Mail senden dürfen. Abmeldung jederzeit per Link in der Mail.

    Picking the model is just the start

    Turning a model into real business value takes clean data, clear processes and the right distribution. That's where we come in.

    Need clean data as input?

    No model in the world will rescue bad input data. We bundle, clean and enrich your data so every token counts.

    Go to DataNaicer

    Want to turn data into content directly?

    Product descriptions, SEO copy and variants at scale — generated from data instead of typed laboriously into ChatGPT.

    Go to ContentNaicer

    Need continuous reach?

    News Stream takes over LinkedIn, blog and newsletter fully automatically — we pick the right AI model per format for you.

    Go to News Stream

    All AI models – pricing in detail

    List prices per 1 M tokens (USD) plus a benefit recommendation per model.

    Data status: Monday, 14 September 2026 at 06:00|Automatic crawl once a week — 6/6 Anbieter, 38 neu, 5 geändert, 14 entfernt · next run: 21/09/2026

    OpenAI

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    GPT-5$5.00$15.00400kWhen you need a real second opinion: strategy papers, legal analysis, hard code refactors, agents that orchestrate multiple tools cleanly.
    Subscription equivalent: ChatGPT Plus 20 $ / Pro 200 $ pro Monat
    GPT-5 mini$0.40$1.60400kThe workhorse for production apps: chat assistants, content drafts, RAG answers, mid-tier tool calling.
    GPT-4o$2.50$10.00128kVoice assistants, image description and OCR — anywhere you need to mix text, image and audio.
    o4-mini$1.10$4.40200kMaths, logic, unit-test generation and coding agents when you need reasoning but don't want to pay GPT-5 prices.
    GPT-6 Astra$10.00$50.00-Nutze mich für die anspruchsvollsten und komplexesten Projekte, die eine umfassende Lösung erfordern.
    GPT-5.6 Sol$4.00$20.00-Ich bin ideal für anspruchsvolle agentische Anwendungen, die intelligente Entscheidungen und komplexe Arbeitsabläufe benötigen.
    GPT-5.6 Terra$2.00$12.00-Verwende mich, wenn du eine ausgewogene Leistung für effiziente und hochvolumige Aufgaben benötigst.
    GPT-5.6 Luna$0.20$1.20-Ich bin perfekt für schnelle, kostengünstige und alltägliche Aufgaben, die keine hohe Komplexität erfordern.

    Anthropic

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Claude Fable 5$10.00$50.001MAnthropic's new top tier (Mythos class). For the truly hard problems: multi-step research agents, autonomous coding sessions running for hours, deep legal analysis.
    Subscription equivalent: Nur über Claude Max 200 $ oder API
    Claude Opus 4.8$5.00$25.001MCurrently the first choice for long-running coding agents (Claude Code, Cursor) and whole codebases or document stacks in context. Writes at a top level with style and consistency.
    Subscription equivalent: Claude Pro 20 $ / Max 100–200 $ pro Monat
    Claude Sonnet 4.6$3.00$15.001MThe workhorse for coding agents and long-context RAG — cheaper than Opus, almost as good for most tasks.
    Claude Haiku 4.5$1.00$5.00200kCustomer-support bots, ticket routing, sentiment and intent classification at high volume.
    Fable 5.1$10.00$50.00200kWenn du intelligente Agenten für komplexe, langlaufende Aufgaben benötigst, ist dieses Modell ideal.
    Opus 5$5.00$25.00200kFür komplexe agentische Codierung und Unternehmensaufgaben, die hohe Präzision erfordern, ist es perfekt.
    Sonnet 5$2.00$10.00200kFalls du ein leistungsstarkes Modell für Coding und Agenten brauchst, das schnell Ergebnisse liefert, ist dieses Modell gut.
    Haiku 4.5$1.00$5.00200kDieses Modell ist ideal, wenn du schnelle und kosteneffiziente Ergebnisse für deine Anfragen suchst.

    Google

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Gemini 2.5 Pro$1.25$10.001MWhen you need to understand entire codebases, videos or hundreds of PDFs at once — the king of long context.
    Subscription equivalent: Gemini Advanced 21,99 $ / AI Ultra 250 $ pro Monat
    Gemini 2.5 Flash$0.35$1.051MRAG, translations, image recognition at scale — excellent price/performance with long context.
    Gemini 2.5 Flash-Lite$0.10$0.401MTagging product data, bulk translations, simple classification — when every cent per call matters.
    Gemini 3.8 Flash$0.75$3.752MDu kannst es für anspruchsvolle Softwareentwicklung und komplexe Unternehmensabläufe verwenden.
    Gemini 3.7 Flash$0.75$3.752MNutze es für alltägliches Coding, den Einsatz von Agenten und zuverlässige mehrstufige Ausführung.
    Gemini 3.6 Flash$0.75$3.752MDieses Modell ist gut für allgemeine agentische und alltägliche Aufgaben geeignet.
    Gemini 3.5 Flash$1.50$9.002MIdeal für schnelle, routinemäßige und hochdurchsatz-Workloads.
    Gemini 3.5 Flash-Lite$0.30$2.502MVerwende es für kostengünstige, hochvolumige Agentenaufgaben, Übersetzungen und einfache Datenverarbeitung.
    Gemini 3.1 Flash-Lite$0.25$1.502MPerfekt für kostengünstige, hochvolumige Agentenaufgaben, Übersetzungen und einfache Datenverarbeitung.
    Gemini 3.1 Pro Preview$2.00$12.002MSetze es für multimodales Verständnis, agentische Fähigkeiten und 'Vibe-Coding' ein.
    Gemini 3 Flash Preview$0.50$3.002MNutze es für grundlegende Geschwindigkeit und Intelligenz in deiner Anwendung.

    Meta (Llama)

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Llama 4 Maverick$0.50$1.501MOpen-weights flagship for on-prem and fine-tuning when you want to host or adapt the weights yourself.
    Llama 4 Scout$0.15$0.6010MLoad huge knowledge bases fully into context without having to build a RAG pipeline.

    Mistral

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Mistral Large 2$2.00$6.00128kFirst choice for GDPR requirements and EU data residency. Multilingual, solid function calling.
    Subscription equivalent: Le Chat Pro 14,99 € pro Monat
    Mistral Small 3$0.20$0.6032kEU-compliant low-latency apps, edge deployments, fast internal tools.
    Mistral Large$0.50$1.50k.A.Du kannst damit komplexe Aufgaben bearbeiten und als Basis für verschiedene KI-Anwendungen verwenden.
    Mistral Small$0.02$0.06k.A.Du kannst diesen Dienst für kostensensible Projekte und alltägliche Aufgaben nutzen, um erste Schritte zu machen.
    Mistral Medium$0.27$0.81k.A.Du kannst diesen Dienst für die meisten Aufgaben, inklusive Coding und Dokumentenanalyse, verwenden.

    xAI

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Grok 4$3.00$15.00256kWhen real-time web and X data have to be part of the answer — e.g. trend research or social monitoring.
    Subscription equivalent: X Premium+ 40 $ / SuperGrok 30–300 $ pro Monat
    Grok 4 mini$0.30$1.50128kFast, cheap answers with up-to-date web knowledge.

    DeepSeek 🇨🇳

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    DeepSeek-V3.2$0.27$1.10128kVery cheap general-purpose model with surprisingly strong coding performance — the price breaker for volume.
    DeepSeek-R1$0.55$2.19128kFrontier-level reasoning at a fraction of the cost of GPT-5 or Opus 4.5.

    Alibaba Qwen 🇨🇳

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Qwen3-Max$1.20$6.0032KTop all-rounder from China, very strong in coding and multilingual tasks (CN, EN, DE).
    Qwen3-Coder$0.30$1.201MSearch huge repos, refactor, generate tests — cheap coding agents.

    Moonshot Kimi 🇨🇳

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Kimi K2$0.60$2.502MLong documents, research agents, "read this book and answer my questions" scenarios.

    Zhipu GLM 🇨🇳

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    GLM-4.6$0.60$2.20200kSolid all-rounder with good tool use and CN/EN strength — happy to use as a backup router target.

    Baidu Ernie 🇨🇳

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Ernie 4.5 Turbo$0.55$2.20128kChinese-language customer communication, marketing content for the CN market.

    ByteDance Doubao 🇨🇳

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    Doubao Pro 1.5$0.11$0.28256kExtremely cheap for mass content, social tagging and content moderation.

    MiniMax 🇨🇳

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    MiniMax M2$0.30$1.201MMultimodal apps (text + image) at a low price, popular in Asia.

    DeepSeek

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    deepseek-flash$0.01$0.011MDu kannst es für schnelle, kostengünstige Aufgaben wie JSON-Ausgabe und Tool-Nutzung verwenden, wenn Geschwindigkeit entscheidend ist.
    deepseek-v4-pro$0.04$0.081MDieses Modell ist ideal für dich, wenn du anspruchsvolle Aufgaben mit komplexer Logik und Tool-Nutzung bearbeiten möchtest.

    Alibaba (Qwen)

    ModelInput / 1MOutput / 1MContextWhat you'd use it for
    qwen3.8-max$2.00$6.001MDu möchtest anspruchsvolle Aufgaben lösen, die hohe Genauigkeit und tiefes Verständnis erfordern.
    qwen3.8-max-0902$2.00$6.001MDu möchtest anspruchsvolle Aufgaben lösen, die hohe Genauigkeit und tiefes Verständnis erfordern.
    qwen3.7-max$2.50$7.501MIdeal für komplexe Problemstellungen, die ein detailliertes Denkvermögen erfordern.
    qwen3.7-max-2026-06-08$2.50$7.501MIdeal für komplexe Problemstellungen, die ein detailliertes Denkvermögen erfordern.
    qwen3.7-max-2026-05-20$2.50$7.501MIdeal für komplexe Problemstellungen, die ein detailliertes Denkvermögen erfordern.
    qwen3.7-max-preview$2.50$7.501MDu möchtest zukünftige Funktionen und verbesserte Denkfähigkeiten testen.
    qwen3.7-max-2026-05-17$2.50$7.501MIdeal für Aufgaben, die ausschließlich den 'Thinking Mode' benötigen.
    qwen3.6-max-preview$1.30$7.80128KDu suchst eine gute Balance aus Kosten und Leistung für typische Aufgaben.
    qwen3-max-2026-01-23$1.20$6.0032KFür allgemeine Textgenerierungsaufgaben, die eine solide Leistung erfordern.
    qwen3-max-2025-09-23$1.20$6.0032KIdeal für direkte Textgenerierung ohne komplexere Denkprozesse.
    qwen3-max-preview$1.20$6.0032KDu möchtest eine solide Grundleistung mit der Option, Denkmodi zu nutzen.
    qwen-max$1.60$6.40No tiered pricingGut geeignet für Standard-Textgenerierungsanfragen im Nicht-Denkmodus.
    qwen3.7-plus$0.40$1.60256KDu suchst eine effiziente Lösung für viele Textgenerierungsaufgaben mit gutem Preis-Leistungs-Verhältnis.
    qwen3.7-plus-2026-05-26$0.40$1.60256KDu suchst eine effiziente Lösung für viele Textgenerierungsaufgaben mit gutem Preis-Leistungs-Verhältnis.
    qwen3.6-plus$0.50$3.00256KDu benötigst ein ausgewogenes Modell für eine Vielzahl von Anwendungen mit angemessenen Kosten.
    qwen3.6-plus-2026-04-02$0.50$3.00256KDu benötigst ein ausgewogenes Modell für eine Vielzahl von Anwendungen mit angemessenen Kosten.
    qwen3.5-plus$0.40$2.40256KDu möchtest schnelle und kosteneffiziente Textgenerierung für Standardanwendungen.
    qwen3.5-plus-2026-04-20$0.40$2.40256KDu möchtest schnelle und kosteneffiziente Textgenerierung für Standardanwendungen.
    qwen3.5-plus-2026-02-15$0.40$2.40256KDu möchtest schnelle und kosteneffiziente Textgenerierung für Standardanwendungen.
    qwen-plus$0.40$1.20256KFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen-plus-latest$0.40$1.20256KFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen-plus-2025-12-01$0.40$1.20256KFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen-plus-2025-09-11$0.40$1.20256KFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen-plus-2025-07-28$0.40$1.20256KFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen-plus-2025-07-14$0.40$1.20No tiered pricingFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen-plus-2025-04-28$0.40$1.20No tiered pricingFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen-plus-2025-01-25$0.40$1.20No tiered pricingFür alltägliche Textaufgaben, bei denen Zuverlässigkeit und Kosten im Vordergrund stehen.
    qwen3.7-plus-us$0.40$1.60256KOptimal, wenn du eine lokalisierte, kosteneffiziente Lösung in den USA suchst.
    qwen-plus-us$0.40$1.20256KIdeal für Standardanwendungen in der US-Region mit Fokus auf Kostenersparnis.
    qwen-plus-2025-12-01-us$0.40$1.20256KIdeal für Standardanwendungen in der US-Region mit Fokus auf Kostenersparnis.
    qwen3.8-flash$0.15$0.471MDu benötigst extrem schnelle Antworten zu sehr niedrigen Kosten für Massenverarbeitung.
    qwen3.7-flash$0.03$0.1332KPerfekt für Anwendungen, die schnelle, kostengünstige Antworten für kürzere Texte benötigen.
    qwen3.7-flash-2026-07-15$0.03$0.1332KPerfekt für Anwendungen, die schnelle, kostengünstige Antworten für kürzere Texte benötigen.
    qwen3.6-flash$0.25$1.50256KDu suchst einen schnellen und kostengünstigen Allrounder für allgemeine Aufgaben.
    qwen3.6-flash-2026-04-16$0.25$1.50256KDu suchst einen schnellen und kostengünstigen Allrounder für allgemeine Aufgaben.
    qwen3.5-flash$0.10$0.401MDu brauchst Top-Geschwindigkeit und die geringsten Kosten für längere Texte.
    qwen3.5-flash-2026-02-23$0.10$0.401MDu brauchst Top-Geschwindigkeit und die geringsten Kosten für längere Texte.
    qwen-flash$0.05$0.40256KDu legst größten Wert auf minimale Kosten und hohe Geschwindigkeit für kürzere Anfragen.
    qwen-flash-2025-07-28$0.05$0.40256KDu legst größten Wert auf minimale Kosten und hohe Geschwindigkeit für kürzere Anfragen.

    Note: Prices are list prices for the providers' APIs (as of September 2026) and change regularly. Discounts via batch, cache or volume contracts and regional hosting surcharges (Azure, Vertex AI, Bedrock) are not included.

    Common questions about AI model costs

    5 things nobody says out loud — and that aren't on any price list.

    • Are output tokens really more expensive than input tokens?

      Yes, usually by a factor of 3–5×. If you generate long answers you pay disproportionately. Trick: keep system prompts short, cap answer length.

    • Is a large context window automatically more expensive?

      Long prompts cost linearly with token count. A 1M context only pays off if you actually need it — RAG is often cheaper.

    • What are reasoning or thinking tokens?

      The GPT-o series, DeepSeek-R1 and Claude Opus “think” internally. These thinking tokens get billed too, which is why reasoning models often come out higher than expected.

    • What does multi-model routing actually give you?

      Route simple tasks to Flash-Lite or DeepSeek and hard ones to Sonnet or GPT-5 — in our projects this typically saves 60–80 % of the cost.

    • Which AI models can I use under GDPR?

      For GDPR-critical workloads, Mistral (EU hosting) or self-hosted Llama are usually the only clean choice. Chinese models have no place in sales or personal data.

    Not sure which model fits your use case?

    We pick for you — data-driven, vendor-neutral and with an eye on total cost of ownership.