AI Tools टूल अपडेट किया गया 4 दिन पहले

AI API Cost Calculator

See the real price of an AI call before you ship. Pick a provider and model from the live models.dev catalog, enter your token usage including cache, reasoning and audio, and get a per-request cost plus a monthly projection. Compare up to 4 models at once, browse and sort the whole catalog by price, and catch context or output-limit mistakes before they cost you money.

हमेशा मुफ़्त उपयोग में आसान तत्काल परिणाम निजी और सुरक्षित

उपयोग करें AI API Cost Calculator

मूल्य निर्धारण डेटा लोड हो रहा है...

मॉडल

उपयोग

अनुरोध पर

त्वरित परिदृश्य

टोकन अनुमानक

अनुमानित टोकन: 0
अनुरोध के अनुसार
$0.0000
प्रति 1M टोकन
$0.00
मासिक
$0.00
टोकन / अनुरोध
0

लागत का विवरण

घटक टोकन दर / 1M लागत
विवरण देखने के लिए एक मॉडल चुनें और उपयोग दर्ज करें।
विस्तार, दरें और चेतावनियाँ गणना के बाद यहाँ दिखाई देंगी।

तुलना करने के लिए मॉडल

साझा उपयोग (प्रति अनुरोध)

तुलना

मॉडल इनपुट / 1M आउटपुट / 1M अनुरोध के अनुसार मासिक
कम से कम दो मॉडल चुनें और साझा उपयोग दर्ज करें, फिर तुलना करें पर क्लिक करें।
मॉडल प्रदाता संदर्भ इनपुट / 1M आउटपुट / 1M कैश रीड / 1M स्थिति उपयोग करें
पूरे कैटलॉग में खोजने के लिए कम से कम एक अक्षर टाइप करें।
टूल लिंक
क्लिपबोर्ड पर कॉपी हो गया।
इस उपकरण को बेहतर बनाने में मदद करें

सामग्री पढ़ें

जानें कि यह उपकरण कैसे काम करता है और इसका उपयोग कब करना है।

What Is the AI API Cost Calculator?

The AI API Cost Calculator is a free online tool that answers one practical question: how much will this AI request cost me? You choose a provider and a model, enter how many tokens your request uses, and the tool returns the price in US dollars for a single request and for a whole month of traffic at your chosen volume.

It is built on the public pricing catalog at models.dev, so the rates come from the providers themselves rather than from guesswork. The catalog covers thousands of models from OpenAI, Anthropic, Google, DeepSeek, Meta, Mistral, Alibaba, xAI, Amazon, Cohere and many others. Pricing is stated the way providers bill it: in USD per 1 million tokens, with separate rates for input, output, cached text, reasoning steps and audio.

A token is the billing unit of every large language model. As a rough guide, one token is about four characters of English text, so a 400-word article is close to 500 tokens. Because input and output are priced differently, and because cache and reasoning have their own rates, a mental calculation is almost always wrong. This tool does it exactly.

Why Use It — and Who It Is For

Model pricing is easy to misread. A rate that looks tiny per thousand tokens becomes a serious bill at production volume, and an apparently expensive model can turn out to be the cheapest once prompt caching is counted. Real people use this calculator for:

  • Developers who need a cost answer before wiring a model into a product.
  • Founders and product managers who are budgeting a monthly AI bill or choosing between providers.
  • Consultants and agencies who must quote a fixed price for a project with unknown traffic.
  • Students and researchers comparing open-weight models for an experiment on a tight budget.
  • Anyone with an unexpected invoice who wants to see where the money went, component by component.

The question people actually need answered is rarely "which model is smartest". It is "which model fits my budget at ten thousand requests a day", and that is what this calculator measures.

How This Tool Works

Everything on the page comes from one pricing catalog. When the page loads, the browser asks our server for the provider list. The server downloads the official models.dev API file, caches it for 24 hours, and also keeps a stored copy on disk, so the tool keeps working even if the upstream site is briefly unreachable. Every response carries the last update time and a stale flag, and the page shows a warning when pricing data could not be loaded at all.

When you change a number, the page waits 350 milliseconds before sending the request, so fast typing does not flood the server. The calculate endpoint then runs these steps on our backend:

  1. Lookup. Your provider and model are matched against the catalog. If either is unknown you get the message "Model not found in catalog."
  2. Validation. At least one token count must be greater than zero, otherwise the server replies "Enter at least one token count greater than zero."
  3. Clamping. Each token field is clamped to a maximum of 50,000,000 tokens, and requests per month is clamped to a range of 1 to 100,000,000, so a typo cannot produce a nonsense result.
  4. Line items. Every component with tokens above zero and a published rate becomes one line: tokens divided by 1,000,000, multiplied by the rate. Components with no tokens, or with no published rate, are skipped rather than guessed.
  5. Totals. Line items are summed into a per-request cost, then multiplied by your monthly request count to produce the monthly projection.
  6. Warnings. The server checks your input side (input plus cache read, cache write, reasoning and audio input) against the model context window, and your output tokens against the model maximum output limit, and warns you when either is exceeded.

Rates are applied with the same fallbacks providers use. Cache write falls back to cache read and then to the input rate when a provider does not publish it. Reasoning tokens use the reasoning rate when available and otherwise the input rate. Audio input and audio output fall back to the standard input and output rates.

Key Features

  • Three working tabs. Calculate Cost, Compare Models and Pricing Browser, each with its own panel.
  • Live pricing. Rates are refreshed from models.dev on a 24-hour cycle, and every response reports the last update time plus a stale flag.
  • Full token breakdown. Input, output, cache read, cache write, reasoning, audio input and audio output are each billed at their own rate and listed in a table with tokens, rate and cost.
  • Four headline numbers. Per request, weighted price per 1M tokens, monthly total and tokens per request.
  • Monthly projection. A slider from 1 to 10,000,000 requests per month recalculates the bill as you drag it.
  • Quick scenario chips. Chat assistant, RAG pipeline, Code generation, Deep reasoning and Voice assistant fill the token fields with realistic starting values in one click.
  • Side-by-side comparison. Put 2 to 4 models against one shared usage profile and get a table plus a monthly-cost bar chart.
  • Pricing browser. Search the catalog, sort by input, output or cache-read price, switch between cheapest first and priciest first, filter to reasoning models only, and load any row into the calculator with its Use button.
  • Token estimator. Paste any text for a quick token count using the four-characters-per-token rule of thumb, which suits English and over-estimates for Chinese or Japanese.
  • Context and output warnings. Automatic alerts when your numbers cannot fit the model you selected.
  • Model meta badges. Provider, status, reasoning, tool calls, open weights, context size and maximum output are shown next to the model you pick, with deprecated models marked in red.

How to Use — Step by Step

  1. Open the Calculate Cost tab, which is selected by default.
  2. Choose a Provider from the dropdown. The provider list loads from the live catalog.
  3. Choose a Model. Type in the filter box above the list to narrow it, then click the model you want. Meta badges appear underneath.
  4. Enter your usage per request: Input tokens (default 1000), Output tokens (default 500), and any of Cache read, Cache write, Reasoning, Audio input and Audio output that apply to your workload.
  5. Drag the Requests per month slider to your real traffic volume, or click a quick scenario chip such as RAG pipeline to fill the fields with sensible defaults.
  6. Read the results: four stat cards, then the cost breakdown table, then the note summarizing per-request cost and monthly total. Warnings appear above the note if your input exceeds the context window or your output exceeds the model limit.
  7. Optional: paste your prompt into the Token Estimator box to check that your token guess is realistic.
  8. Switch to the Compare Models tab. Two rows are already there; click Add for a third and fourth. Set the shared usage numbers, then click Compare to get the table and bar chart.
  9. Switch to the Pricing Browser tab, type at least one character (for example gpt, claude or llama), pick a sort field and direction, and optionally enable Reasoning models only. Click Use on any row to send that model back to the calculator.

Limits and Rules

  • Pricing endpoints are rate limited to 120 requests per minute. Normal interactive use never comes close to this.
  • Each token field accepts up to 50,000,000 tokens; larger values are silently clamped to that ceiling.
  • Requests per month is clamped between 1 and 100,000,000.
  • At least one token count must be above zero, or the server refuses to calculate.
  • Comparison accepts at least 2 and at most 4 models per run.
  • The pricing browser returns the top 200 matches, sorted by the field you choose.
  • All prices are USD per 1 million tokens, taken from published list prices. Tiered volume discounts, batch-processing discounts and negotiated enterprise rates are not modelled.
  • The pricing catalog refreshes every 24 hours. If a provider changed its rates today, the page may show yesterday rates until the cache expires, and the API marks that snapshot with a stale flag.
  • No account, no API key and no captcha are needed. The calculator only reads pricing data; it never calls a model on your behalf.

When to Use It — and When Not To

Use it when you are picking a model for a new feature, checking whether prompt caching will pay for itself, forecasting next month spend, justifying a provider switch to your team, or explaining an invoice line to a client. It is also useful for spotting models whose output price is several times their input price, which matters a lot when your app generates long answers.

Do not rely on it when you need an exact invoice. The calculator uses published list rates only: batch discounts, long-context surcharges, cached-token promotions and per-seat subscriptions are not included. It also cannot price features the catalog does not publish rates for, in which case that line is simply omitted. For a final contract or a formal budget, confirm the numbers against the provider own pricing page.

Frequently Asked Questions

अक्सर पूछे जाने वाले प्रश्न

परिणाम पर भरोसा करने से पहले उपयोगकर्ताओं द्वारा पूछे जाने वाले प्रश्नों के संक्षिप्त उत्तर।

Yes. It is free to use, needs no account, no API key and no captcha. The pricing endpoints allow 120 requests per minute, which is far more than interactive use requires.

From the public models.dev catalog. Our server downloads the catalog, caches it for 24 hours and stores a local copy, so the rates you see are the providers published rates, refreshed automatically.

The estimator uses a simple rule of about four characters per token, which works well for English and over-estimates for Chinese and Japanese, where a single character can be one or two tokens. For a final number, use the token counter of the provider you plan to use.

Either you entered zero tokens for it, or the provider does not publish a rate for that component. The calculator never invents a price: a missing rate stays missing instead of being silently substituted.

Your input side (input plus cache read, cache write, reasoning and audio input) is larger than the model context window, or your requested output is above the model maximum output. The model would reject that request in production, so fix the numbers before budgeting from them.

One comparison run accepts between 2 and 4 models. Run a second comparison for the rest, keeping the shared usage numbers the same so the monthly totals stay comparable.

No. It only reads pricing data and multiplies it by your inputs. Nothing you type is sent to a model provider, and no token is ever billed to you.

टिप्पणियाँ (0)

उपयोगकर्ता चर्चा, असामान्य परिस्थितियाँ, और अनुवर्ती सुझावों के लिए।
टिप्पणी करें
आपकी टिप्पणी सबमिट करने के बाद सार्वजनिक रूप से दिखाई देगी।
अभी तक कोई टिप्पणी नहीं। टिप्पणी करने वाले पहले व्यक्ति बनें!
हमें बताएं!
इस उपकरण में समस्या की रिपोर्ट करें

आपको जो समस्या आई, उसका वर्णन करें ताकि हम इसकी जांच कर सकें और टूल में सुधार कर सकें।

सबसे ज्यादा मदद क्या करता है
इनपुट, अपेक्षित परिणाम, वास्तविक परिणाम, ब्राउज़र/डिवाइस, और जब सहायक हो तो एक स्क्रीनशॉट शामिल करें।
स्क्रीनशॉट बटन ब्राउज़र में पृष्ठ को कैप्चर करता है और उत्पन्न छवि को इस फॉर्म से संलग्न करता है।
पूरे पेज का स्क्रीनशॉट
आपके ब्राउज़र में newisty के साथ कैप्चर किया गया।
अभी तक कोई स्क्रीनशॉट नहीं लिया गया है।
उपकरण रिपोर्ट