AI Companion App Token Costs Explained
AI companion app token costs are easy to miss because they usually sit underneath the subscription price.
The pricing page may show a subscription first. Then, after you start using voice, images, videos, private packs, or other premium features, the app may ask you to buy more credits. That does not always mean the app is doing anything shady. It does mean the advertised subscription is not always the same as the realistic monthly cost.
Here is the plain version: tokens are best understood as paid credits for expensive features. Text chat may be included in the subscription, while media generation, voice, video, or premium content spend from a separate balance. If you use those features often, the real bill can rise quickly.
Key Takeaways
- Tokens are app credits, not the same thing as technical AI model tokens
- The biggest token costs usually come from images, videos, voice, and private content
- Candy AI’s terms say the monthly allowance is disclosed at checkout and is generally 100 tokens; do not assume that amount applies to every offer or region
- Refunds, cancellation, and account deletion are separate steps, and the original billing channel often controls them
- The safest move is to test monthly before committing to an annual plan
Evidence status (checked August 15, 2026): The app-specific rows below are desk-researched from official terms, help centers and support documentation. We did not complete a logged-in checkout or token-usage test for this update. Exact regional plan prices, package sizes and cost-per-action figures are therefore left unverified rather than estimated. Local-AI sections use official SillyTavern, Layla and Backyard AI documentation; hardware economics are decision guidance, not a benchmark promise.
What AI companion app tokens actually mean
In normal AI engineering, a token is a chunk of text processed by a model. In consumer AI companion apps, the word usually means something different: a credit balance inside the app.
Think of tokens like arcade credits. Your subscription opens the door, but certain machines still cost credits to play. Depending on the app, tokens may be used for:
- AI images
- AI videos
- Voice messages or calls
- Premium model responses
- Private content packs
- Extra companion slots
- Higher daily limits
That distinction matters because two apps with similar monthly prices can have very different real costs. A $14 plan with unlimited text and no credit layer can be cheaper than a $10 plan that constantly asks you to top up tokens.
The best way to compare is not “Which app has the lowest monthly price?” It is “Which app includes the features I will actually use?”
Verified token, cancellation, and refund snapshot
This table focuses on what official documentation currently supports. A blank or “checkout required” field is more useful than a guessed number.
| App | Free tier or included allowance | Token or add-on rule | Cancellation and refund friction | Evidence status |
|---|---|---|---|---|
| Candy AI | Official terms describe a general five-message maximum for free users. Paid messaging is described as unlimited, with a checkout-disclosed monthly token allowance generally stated as 100. | Subscription tokens can fund extended features such as images or voice notes. Remaining subscription tokens expire at period end and do not carry over or get refunded after cancellation. | In-account cancellation; access continues to period end. Official refund help generally limits eligible card-payment requests to 24 hours and denies them after more than 20 tokens are used. | Official terms/help checked; exact current price and checkout allowance not verified |
| CrushOn AI | Exact current free limits and plan prices require checkout verification. | Purchased Diamonds are non-refundable. Activated chat packages and digital products are non-refundable. | Cancel through Apple, Google, or the original payment platform. Most started subscription periods are non-refundable; deletion may be blocked by an active subscription. | Official FAQ checked; checkout limits not verified |
| SpicyChat | Exact current free limits, plan prices, and token/context mechanics require checkout verification. | Payment route matters: official support distinguishes recurring web-card billing from one-time Pay By Bank and manually linked Boosty or crypto payments. | Website card subscriptions use the account portal. The refund policy excludes partial-month credits and notes that some payment methods cannot be refunded; Apple handles Apple purchases. | Official support/refund pages checked; checkout mechanics not verified |
| Replika | Exact current plan prices and benefits require checkout or app-store verification. | No verified token allowance is published here; do not treat Replika as a token-priced app without current evidence. | Digital subscriptions are described as used immediately and non-refundable. Apple and Google handle store refunds. Deleting the app or account does not cancel billing. | Official help checked; current price not verified |
| Nomi AI | Exact current plan prices require checkout or app-store verification. | No verified token allowance is published here; compare it as a subscription product unless current official evidence says otherwise. | Cancel through the original purchase platform before renewal. Benefits continue to period end; deleting the app does not cancel the subscription. Payments are generally non-refundable. | Official refund/cancellation policies checked; current price not verified |
The first limiter matters more than a “free” badge. It may be a message cap, timer, energy balance, media-credit allowance, ad load, or premium-only feature. If the official page does not state the limiter, treat it as unknown until you test the account yourself.
Why token pricing exists
Token pricing exists because some companion features are more expensive to run than plain text chat.
Text messages are comparatively cheap. Image generation, video generation, real-time voice, high-quality voices, and long context windows cost more infrastructure. A token system lets the company charge heavy users more without raising the base subscription for everyone.
That can be fair. A user who sends a few messages a week should not necessarily pay the same as a user generating dozens of images and videos. The problem is transparency. Many users subscribe for the emotional or creative experience, then discover the credit layer after they are already attached to a companion or story.
Before paying, look for three things:
- What the base subscription includes
- Which features spend tokens
- How many tokens a normal week would use for you
If the pricing page does not answer those clearly, assume the subscription price is only a floor.
Where token costs usually show up
The most common token-gated features are media and voice.
Images and videos. Visual companion apps often include a small monthly allowance, then charge more once that allowance runs out. This is the easiest place to overspend because each generation feels small in the moment.
Voice. Some apps include voice in the paid plan, some make it free, and some route higher-quality voice through credits. If voice is your main use case, verify this before subscribing.
Private packs or premium scenes. Some adult or roleplay-focused apps sell extra content packs with tokens. These can make the app feel cheaper upfront and more expensive in practice.
Higher model access. A few apps use tiers or credits to gate longer context, smarter responses, faster replies, or priority access. This is not always called “tokens,” but it has the same budget effect.
Candy AI is the clearest documented example in this guide. Its official terms say paid subscribers receive a monthly token amount disclosed at checkout, generally 100 tokens, and that extended features such as image generation or voice notes can use tokens. Because the allowance is checkout-disclosed and can vary, verify your own offer rather than budgeting from a review screenshot.
That is not a reason to avoid Candy AI by itself. It is a reason to budget for more than the headline plan if you care about generated media.
Read our deeper Candy AI pricing breakdown for the app-specific version.
AI companion token costs by app type
Here is a practical way to compare the category.
| App type | Pricing pattern | Token risk | Best fit |
|---|---|---|---|
| Flat subscription companion | Monthly or annual plan covers most use | Low | Text chat, memory, long-term companion use |
| Media-first companion | Subscription plus image/video credits | High | Users who mainly want visuals |
| Free-first character platform | Free core use, optional paid tier | Low to medium | Casual roleplay, broad character browsing |
| Tiered power-user app | Higher plans unlock memory, context, or advanced tools | Medium | Heavy users who know they will use the extras |
The three different token bills buyers confuse
The word “token” now describes three different meters. Mixing them produces bad comparisons.
| Meter | What it measures | Where you see it | What changes the bill |
|---|---|---|---|
| App credits / neurons / coins | Product actions | Candy AI, Joi AI and other media-led apps | Images, video, voice, romantic messages, gifts or premium scenes |
| Model API tokens | Text or multimodal input and output processed by a hosted model | BYOK frontends such as Janitor AI or SillyTavern connected to a cloud API | Model price, context size, response length, retries and summaries |
| Local compute | Hardware work, not a token wallet | SillyTavern or Layla using an on-device model | Model size, RAM/VRAM, generation time, electricity and upgrades |
Joi AI is an especially clear warning against stopping at the subscription price. Its current official terms say Premium may coexist with separate bot subscriptions, while “neurons” fund actions outside those subscriptions. The terms list four neurons to read one romantic message, up to 120 for a photo and up to 1,900 for a video, with regional and promotional variation. That means a video can consume the equivalent of hundreds of romantic-message actions. Always verify the live wallet price before converting those units into dollars.
Hosted subscription vs BYOK API vs local AI
| Route | How you pay | First limiter | Privacy boundary | Best fit |
|---|---|---|---|---|
| Hosted companion | Monthly/annual plan, sometimes plus credits | Message cap, media wallet, model tier or feature gate | Companion company and its processors | People who want polish and low setup |
| BYOK hosted model | Per-input/output model tokens, sometimes plus frontend costs | Spend cap, context cost or rate limit | Frontend, API provider and any proxy | Power users who want model choice |
| Fully local model | Hardware, power, storage and time | RAM/VRAM, speed and model quality | Your device and backups | Privacy/control users with setup tolerance |
SillyTavern does not remove model cost by itself. Its official docs describe it as a locally installed interface without inference capabilities. Connect a hosted API and you inherit that provider’s token price and privacy policy. Connect a local llama.cpp-style backend and the per-message invoice disappears, but the compute bill moves onto your device.
Layla packages this distinction more clearly for phones. Its official site says most modes use on-device models and that optional Cloud mode sends requests to the selected hosted provider. The same app can therefore sit in two cost columns depending on the mode.
How to estimate a BYOK API month
Hosted model pricing normally charges separately for input and output. Use this structure rather than guessing from the price per million tokens:
Monthly API cost =
(input tokens × input rate)
+ (output tokens × output rate)
+ image / voice / embedding charges
+ retries and regenerated replies
Long-running companionship can make input grow because the frontend resends character instructions, recent history, retrieved memories and lore with every turn. A 200-token reply can require thousands of input tokens. Context caching or summaries can reduce cost, but only if the provider and frontend use them effectively.
For one week, record the provider dashboard total and divide it by completed conversations—not messages. Regenerations, alternate replies and abandoned calls are real spend. Set a hard provider budget before connecting the key to any third-party frontend.
How to estimate a local AI month
Local inference has no universal price per reply. Use a total-cost view:
Monthly local cost =
hardware amortization
+ electricity
+ storage / backup
+ optional voice or image tools
+ the value of setup and maintenance time
If you already own suitable hardware, the marginal cost can be low. If you buy a GPU only to avoid a modest subscription, the payback can take years. Smaller on-device models may also be less capable than premium hosted models, so privacy and control—not savings alone—are the stronger reason to go local.
The official SillyTavern docs recommend at least 6GB of VRAM for local inference, while noting that the interface itself has minimal requirements. Layla supports phone-focused local model formats and allows users to load GGUF models. These are compatibility claims, not promises that every device will run every model quickly.
Compare the first limiter, not the word “unlimited”
Every route has a first constraint:
- media-led app: credit balance;
- flat companion: premium feature or fair-use policy;
- BYOK API: budget or context-window spend;
- local model: memory, thermal limits or generation speed;
- mobile local model: battery, storage and smaller model capability.
Write that limiter beside the monthly price. It is often more predictive of satisfaction than the advertised plan name.
This is why the “cheapest” app on paper is not always the cheapest app for your actual habits.
If you mainly want text companionship and memory, compare flatter options first. Our Kindroid review and Nomi AI review are good starting points because both apps are stronger fits for users who care about continuity more than paid media generation.
If you mostly want visual roleplay, Candy AI can still make sense, but the token budget is part of the product. Treat it like buying an image allowance, not just a chat subscription.
If you want a large character ecosystem, Character.AI has a different cost profile. Read the Character.AI subscription guide for current free-versus-paid questions; paying should not be assumed to remove filters, age assurance, or other product limits unless official documentation says so.
A simple formula for estimating monthly cost
Use this before you pay:
Real monthly cost = subscription price + expected token top-ups + paid add-ons
Then ask:
- How many images or videos will I generate each week?
- Does voice cost tokens or come with the plan?
- Do tokens reset monthly?
- Do unused tokens roll over?
- What happens when I run out?
- Can I still chat without buying more?
- Is cancellation handled on the web, in the app, or through Apple/Google?
Candy AI’s help center says regular chat remains accessible even without extra tokens, while premium features may require more. That is the kind of detail you want before you subscribe, because it tells you whether running out of tokens pauses the whole app or only the extras.
For an app with a token balance, do a one-week trial math exercise. If you use 25% of the monthly allowance in two days, the base plan is probably not your real monthly cost.
The biggest mistake: annual billing too early
Annual plans are tempting because the displayed monthly equivalent can look much cheaper. But ordinary plan prices and renewal terms in this category can vary by region, platform, promotion, and checkout eligibility.
That discount only helps if your usage matches the included allowance. If you still need regular top-ups, the annual discount may save less than it appears. Worse, you have committed before learning whether you like the app after the novelty period.
The safer path:
- Start with monthly billing.
- Track token use for 30 days.
- Decide whether the app is still useful without extra purchases.
- Only then consider quarterly or annual billing.
This is especially important for companion apps because attachment can distort spending. A small top-up feels easier when it is attached to a character, relationship, or story you have already built.
Which users should avoid token-heavy apps?
Token-heavy pricing is a poor fit if you want predictable cost, private journaling, or emotional-support style companionship. In those cases, you are usually better off with a product where memory, text, and basic companion behavior sit inside the plan.
Be cautious if:
- You dislike tracking credits
- You know you will use images or videos daily
- You are subscribing for emotional support during a vulnerable period
- You are tempted by annual discounts before testing the app
- The pricing page does not explain token spend clearly
Token pricing is not automatically bad. It is just a worse match for users who need a stable budget.
Bottom line
AI companion app token costs matter because they change the real price of the app. The monthly plan is only one part of the budget. The token layer determines what happens when you use the features that make the app feel alive: images, videos, voice, private packs, and higher-limit interactions.
If you want predictable spend, favor apps with flatter pricing and fewer credit prompts. If you want media-heavy companionship, build token top-ups into the monthly cost from the start.
The best rule is simple: do not buy an annual AI companion plan until you know your token burn rate.
If the alternative is a local stack, do the same 30-day exercise with electricity, setup time and actual response quality. The researched private and offline AI companion guide explains why a local interface connected to a cloud API still belongs in the BYOK cost column.
Sources & references
- Candy AI: Terms of Service and refund eligibility help
- CrushOn AI: official FAQ
- SpicyChat: support documentation and refund policy
- Replika: refund policy, cancellation help, and billing after deletion
- Nomi AI: refund policy and cancellation policy
- Joi AI: official terms for Premium, bot subscriptions and neuron costs
- SillyTavern: official documentation and FAQ
- Layla: official on-device and cloud-mode documentation
- Backyard AI: official local desktop architecture post (dated February 2024; verify current product availability)
- CompanionScout pricing guide: The real cost of AI companion subscriptions
- CompanionScout cancellation guide: How to cancel or request a refund from an AI companion app
Frequently asked questions
What are tokens in an AI companion app?
Tokens are app-specific credits used for premium actions such as image generation, video generation, voice features, private content packs, or higher-limit chat features. They are separate from the language-model tokens developers talk about.
Do all AI companion apps charge tokens?
No. Some apps are closer to flat subscriptions, while others combine a monthly plan with a token or credit balance. Candy AI is the clearest example of a subscription plus token model.
Are AI companion app tokens worth buying?
Only if the token-gated features are central to how you use the app. If you mostly want text chat and memory, a flatter subscription is usually easier to budget.
How can I avoid surprise token costs?
Check what your monthly plan includes, which actions spend tokens, when the balance resets, and whether unused tokens roll over. Use monthly billing until you know your real usage.
Do unused AI companion tokens expire?
It depends on the app and credit type. Candy AI's official terms say remaining subscription tokens expire at the end of the subscription period and do not carry over or get refunded after cancellation. Check current official terms before buying credits.
Does deleting an AI companion account cancel the subscription?
Not necessarily. Replika explicitly says deleting the app or account does not cancel billing, and CrushOn AI says account deletion may be blocked while a subscription is active. Cancel through the original billing channel before deleting an account.
Can AI companion tokens or add-ons be refunded?
Refund rules vary. Candy AI limits eligible requests by time and token use, while CrushOn AI says purchased diamonds and activated chat packages are non-refundable. Apple, Google, and other payment platforms may handle their own refund requests.
Is a local AI companion cheaper than app tokens?
Not automatically. Local models avoid per-message fees but move the cost to hardware, electricity, storage, setup and maintenance. They can be economical for heavy use on hardware you already own, while a cloud API may be cheaper for occasional chat.
Does SillyTavern charge tokens?
SillyTavern itself is free and open source, but it is only a frontend. A hosted API connected to it charges according to that provider’s model-token or media pricing; a local model uses your own hardware instead.