17/Services

LLM inference
and GPU services.

Managed inference through an OpenAI-compatible API, dedicated GPU rental and enterprise capacity. One platform, three ways to run.

MetaLlama
QwenQwen
DeepSeekDeepSeek
MistralMistral

Model-family marks identify their respective owners. Availability depends on your deployment; no partnership or endorsement is implied.

OpenAI-compatible REST API
Drop-in replacement with one-line migration
Streaming responses and function calling
Embeddings, reranking and speech-to-text
Auto-scaling with reserved capacity options
Billing in Turkish lira with e-Fatura
01

Managed Inference

Pay per token for open-weight models behind one API.

Shared Pod · Performance Pod
Includes
  • OpenAI-compatible REST API
  • Model families by deployment
  • Streaming
  • Function and tool calling
  • Embeddings, reranking, speech-to-text
From
$0.12 – $2.50
per 1M tokens · by context
Plan Deployment →
02

GPU Rental

Dedicated GPUs for fine-tuning and custom serving.

Performance Pod · Dedicated Pod
Includes
  • Dedicated Blackwell or Ampere GPUs
  • Fine-tuning
  • Custom model serving
  • Private endpoint
From
$0.79 – $2.50
per GPU-hour
Request Capacity →
03

Enterprise

Reserved capacity under a custom annual contract.

Enterprise Pod
Includes
  • Dedicated capacity
  • SLA by agreement
  • KVKK documentation
  • TRY invoicing
Billing
Annual contract
Custom capacity and terms
Explore Enterprise →
18/Model families

Open models,
ready to call.

Choose from open-weight families, Turkish-tuned models and utility models for retrieval and speech. Bring your own fine-tune when you need one.

Llama

  • 3.3 70B Instruct
  • 3.1 8B Instruct
  • Guard 3 8B

Qwen

  • 2.5 72B Instruct
  • 2.5 Coder 32B
  • 2.5 7B Instruct

DeepSeek

  • R1 Distill 70B
  • R1 Distill 32B

Mistral

  • Large
  • Mixtral 8x7B Instruct
  • 7B Instruct

Turkish

  • Tuned chat models
  • Embedding models
  • Custom fine-tunes

Utility

  • Text embeddings
  • Reranking
  • Speech-to-text
Open the model explorer →
19/Illustrative plans and rates

Plan your usage.
Size your capacity.

Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement.

Trial

  • 1M free tokens
  • 5 requests per second
  • No SLA

Starter

  • Pay as you go
  • 20 requests per second
  • Shared pool

Growth

  • $500 per month minimum
  • 20% discount on rates
  • 100 requests per second
  • SLA by agreement
USD per 1M tokens
ContextStarterGrowth
4K$0.15$0.12
8K$0.25$0.20
32K$1.00$0.80
128K—$2.50

Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement.

Choose the capacity your workload needs.

01

Start with an API

For prototypes, assistants and variable traffic: start with managed inference, then measure usage before reserving capacity.

02

Reserve for steady demand

For production traffic with a known baseline: discuss reserved capacity and scheduling against your own latency targets.

03

Isolate for sensitive workloads

For custom weights or stricter boundaries: define dedicated resources, private endpoints and data-handling terms in the deployment agreement.

How to evaluate an LLM inference provider

01

Measure the full response path

Test the same model, prompt length and output limit from your users’ region. Record time to first token, total response time and p95 latency at your expected concurrency. Network proximity alone does not predict model throughput.

02

Map every data transfer

Document prompts, outputs, logs, backups and support access. Ask which parties process them, where they are stored and how long they are retained. Hosting in Türkiye is one input to a KVKK assessment, not a compliance guarantee.

03

Compare cost at the same workload

Compare token-based and GPU-hour costs using expected traffic, idle capacity, storage and operational work. Confirm the billing currency, taxes and any exchange-rate indexing before making a commitment.

Compare the operating model. →

Prepare your first deployment

Bring a representative prompt set, target region, expected requests per minute, peak concurrency and context/output limits. Identify sensitive data and retention requirements. Use the model explorer and pricing guide to define the scope before agreeing on capacity and commercial terms.

21/Performance and trust

Fast where
your data lives.

Latency and throughput depend on the model, input and load. Test them on your workload; availability targets and remedies are defined in the service agreement.

TTFTTime to first tokenMeasure on your workload
TPSTokens per secondMeasure on your workload
SLASLA by agreement
Assess KVKK processing and cross-border transfers
Define permitted data use in your agreement
Confirm retention for prompts, outputs and logs
Agree data-processing roles and responsibilities
Region-specific legal agreements by jurisdiction
Company, privacy & terms
22/Coverage

Real GPUs,
real location.

Every GPU runs in Türkiye today. Other territories are served over the network, and new sites are planned.

Live

Türkiye

RTX PRO 6000 Blackwell clusters with 1.3 TB+ of GPU memory. KVKK-oriented data residency and TRY billing.

Served

MENA and Europe

Served over the network from Türkiye. All GPUs remain in Türkiye.

Planned

Frankfurt · Canada

Compute is hosted in Türkiye today. Frankfurt and Canada are planned locations, not currently available deployment regions.

Questions before your first deployment.

Where does my model run?

Current compute is in Türkiye. Frankfurt and Canada are planned; requests from other markets can be served from Türkiye without moving the GPUs.

Is local compute always faster?

A shorter network route can reduce transport delay, but first-token time also depends on the model, queue and input size. Compare the same model and prompt from your actual users' location.

Does Türkiye hosting automatically mean KVKK compliance?

No. You still need to assess processing grounds, notices, retention, access and any transfers through integrations or support. Regional compute is one part of that assessment.

Does a TRY invoice remove all currency risk?

Not if the agreed price is indexed to USD or another currency. Confirm whether rates are fixed in TRY, which conversion date applies and whether taxes are included.

Company, privacy & terms →

Prepare your
first deployment.

Bring a representative prompt set, target region, expected requests per minute, peak concurrency and context/output limits. Identify sensitive data and retention requirements. Use the model explorer and pricing guide to define the scope before agreeing on capacity and commercial terms.

Plan Deployment → Explore Enterprise
Istanbul, Türkiye · OpenAI compatible · KVKK-oriented
← Pricing Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement.