12/Pricing

GPU and inference
pricing explained.

Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement.

01

Managed Inference

Usage-based model access.

Shared Pod · Performance Pod
Best for
  • APIs
  • Variable demand
  • Experimentation
Billing
$0.12 – $2.50
per 1M tokens · by context
Plan Deployment →
02

GPU Pod

Dedicated or reserved compute.

Performance Pod · Dedicated Pod
Best for
  • Predictable workloads
  • Custom models
  • High throughput
Billing
$0.79 – $2.50
per GPU-hour · Blackwell or Ampere
Request Capacity →
03

Enterprise

Private infrastructure.

Enterprise Pod
Best for
  • Security
  • Regional requirements
  • Large-scale capacity
  • Custom SLA
Billing
Annual contract
SLA by agreement
Explore Enterprise →
In Türkiye, where applicable TRY billinge-Fatura / e-ArşivLocal commercial relationshipPlans, rates and use cases →

What determines your inference bill?

Token-based inference

Estimate input and output tokens separately, including conversation history and retrieved documents. If rates differ, calculate each part at its own rate. Multiply by expected requests; add any separately billed embedding, reranking or audio usage.

Reserved or dedicated GPU

Estimate GPU count × billable hours × agreed hourly rate. Include idle capacity, replicas, storage and network charges where applicable. Compare utilization and operational responsibility with a managed token-based service.

Before approving a quote

Confirm the model/version, context limit, concurrency, billing currency, tax treatment, minimum commitment and overage rules. Specify support hours, retention, SLA measurement and exclusions. A displayed example is not a binding capacity or service commitment.

Comparison

Compare the operating model.

Different tools solve different problems. Compare deployment control, data placement and operating responsibility before comparing a headline price.

Compare the operating model.
CapabilityLLMPodsOpenAI APIRunpodSelf-hosted
Model accessOpen-weight catalog; custom deployments by agreementOpenAI API model catalogCustom GPU workloads and public model endpointsYour choice of weights and serving stack
Data placementTürkiye today; Frankfurt and Canada plannedProject-level residency options; eligibility and scope applyDepends on the selected deployment and productYour infrastructure and connected services
Operating responsibilityManaged inference; isolation options by agreementProvider-operated API servingManaged endpoints or customer-operated GPU containersYou maintain capacity, updates and availability
Commercial modelUsage or reserved capacity; TRY invoicing where agreedUsage-based API pricingCredit-based billing; compute and storage chargesHardware or cloud spend plus your operations cost

Scroll sideways to compare providers →

Feature summary, not a performance benchmark. Options vary by plan and configuration. Sources checked 5 October 2026.

FX risk on every invoice.

If usage is priced in a foreign currency, the same workload can cost more in TRY when exchange rates move. Ask for the billing currency and conversion rule before you commit.

TRY billing is available where applicable; USD-indexed rates may still carry FX exposure.

Same usage. Different local-currency spend.
USD/TRY 30
3 000 TRY
USD/TRY 40
4 000 TRY
USD/TRY 50
5 000 TRY

Illustration only: $100 of usage at hypothetical USD/TRY rates of 30, 40 and 50. Excludes taxes and bank fees; these are not current exchange rates.

14/Enterprise

When one Pod
isn't enough.

Reserve the infrastructure underneath your AI workloads.

01 · Single Pod

One model, one GPU allocation.

02 · Multiple Pods

Replicas that share routing.

03 · Rack

Pods locked into one chassis.

04 · Cluster

Racks behind a private boundary.

  • Dedicated GPU clusters
  • Private networking
  • Reserved capacity
  • Regional architecture
  • Custom model serving
  • Technical account support
  • Enterprise SLA
  • Custom agreements
15/Terminal

Operate Pods
like infrastructure.

A command-line view of the same system. Conceptual UI: commands shown are illustrative until the CLI ships.

16/Documentation

You already know
the API.

POST https://api.llmpods.com/v1/chat/completions
RequestJSON
{
  "model": "llama-70b-instruct",
  "stream": true,
  "messages": [
    { "role": "user", "content": "Hello, Pod." }
  ]
}
Response200 · JSON
{
  "id": "chatcmpl-7f39a21",
  "object": "chat.completion",
  "model": "llama-70b-instruct",
  "choices": [
    { "index": 0,
      "message": { "role": "assistant", "content": "Hello." },
      "finish_reason": "stop" }
  ],
  "usage": { "total_tokens": 15 }
}
Streaming tokensSSE
data: {"delta": {"content": "Hel"}}
data: {"delta": {"content": "lo"}}
data: {"delta": {"content": ","}}
data: {"delta": {"content": " Pod"}}
data: [DONE]
Read Documentation → API Reference →

Give your AI
somewhere better to run.

Plan Deployment → Explore Enterprise
  • OpenAI Compatible
  • Regional Compute
  • Dedicated GPU
  • Managed Infrastructure