What does it actually cost to run your own AI API key?

Everyone asks this before they plug in a key, and almost nobody does the actual math. Here it is, in dollars, at list prices — ~2,000 tokens in, ~500 out, roughly one typical question and answer.

Competitor pricing below is quoted from each company’s own public pricing page and was last checked on 29 July 2026. The per-model costs in the table below assume roughly 2,000 input tokens and 500 output tokens per answer, and are based on each provider’s own published list prices, also checked on 29 July 2026. Prices change — check theirs before deciding.

What $5 buys, model by model

Same $5. Same size of question. Wildly different number of answers, because the price difference between models dwarfs everything else you might optimize.

ModelAnswers on $5Cost per answer
Gemini 3.5 Flash-Lite~2,700~$0.0019
Claude Haiku 4.5~1,100~$0.0045
Claude Sonnet 5~370~$0.0135
Claude Opus 5~220~$0.0227

Look at the spread: Opus 5 costs roughly 12x more per answer than Flash-Lite. That gap is bigger than any difference you’ll find between two subscription plans, two apps, or two pricing pages. The model you pick matters more than almost any other decision you’ll make about cost. If you’re watching spend, the model dropdown is the first place to look — not the app’s price tag.

And Gemini’s Flash line has a genuine free tier, so it’s entirely possible to run a lot of everyday questions through GhostNote without paying an AI provider anything at all.

Why bring-your-own-key is cheap for most people

GhostNote doesn’t resell tokens. You connect your own Anthropic, OpenAI, or Google key — or run a model locally through Ollama, which means no key, no network request, and nothing leaving your machine (see running it fully offline). The app is $11.99/mo or $250 once; whatever you spend on the model itself is separate, and it’s usually a rounding error.

A person using this for a handful of meetings or study sessions a week, on a Flash-Lite or Haiku-class model, is realistically spending cents a month on tokens. At that volume, paying per token and paying a flat subscription come out about the same in your favor either way — except with your own key, you’re never paying for headroom you don’t use.

Where the math flips

Be honest with yourself about volume. If you’re running an expensive model like Opus 5 constantly, all day, every day — not a question here and there, but continuous heavy use — the per-token bill adds up, and a flat-rate subscription somewhere else may end up cheaper for that specific workload than paying list-price API rates. That’s a real tradeoff, not a reason to avoid bring-your-own-key altogether: it just means matching the model to the job. Drop to Haiku or Flash-Lite for routine questions and save Opus-class reasoning for the few answers that actually need it, and the math tends to swing back in your favor.

Rule of thumb: pick the cheapest model that gets the answer right, and only reach for the expensive one when the question is genuinely hard. That one habit affects your bill more than any plan you could switch to.

How this compares to the alternative: paying for tokens through the app

Some tools bundle token cost into a much higher monthly price instead of letting you bring a key. Cluely only becomes invisible to screen sharing at $149.99/mo. Interview Coder charges $299/mo (or $799 lifetime) and caps you at 1,000 credits a month regardless of which model you’d rather use. Ofradr is Windows-only and its own EULA says screen data transmission can’t be disabled without ceasing use of the software. None of the three let you see, per model, what an individual answer actually costs — you’re paying their markup on top, whatever it is.

With your own key you see the real number, at the provider’s real price, every time. That transparency is worth something even before you get to the total.

Full breakdown of what each plan includes is on the pricing page, and a side-by-side of all the tools mentioned here is on the comparison page. If you want the numbers on your own typical week, the usage page walks through it.

See pricing