Start with an API
For prototypes, assistants and variable traffic: start with managed inference, then measure usage before reserving capacity.
Managed inference through an OpenAI-compatible API, dedicated GPU rental and enterprise capacity. One platform, three ways to run.
Model-family marks identify their respective owners. Availability depends on your deployment; no partnership or endorsement is implied.
Pay per token for open-weight models behind one API.
Shared Pod · Performance PodDedicated GPUs for fine-tuning and custom serving.
Performance Pod · Dedicated PodReserved capacity under a custom annual contract.
Enterprise PodChoose from open-weight families, Turkish-tuned models and utility models for retrieval and speech. Bring your own fine-tune when you need one.
Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement.
Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement.
For prototypes, assistants and variable traffic: start with managed inference, then measure usage before reserving capacity.
For production traffic with a known baseline: discuss reserved capacity and scheduling against your own latency targets.
For custom weights or stricter boundaries: define dedicated resources, private endpoints and data-handling terms in the deployment agreement.
Test the same model, prompt length and output limit from your users’ region. Record time to first token, total response time and p95 latency at your expected concurrency. Network proximity alone does not predict model throughput.
Document prompts, outputs, logs, backups and support access. Ask which parties process them, where they are stored and how long they are retained. Hosting in Türkiye is one input to a KVKK assessment, not a compliance guarantee.
Compare token-based and GPU-hour costs using expected traffic, idle capacity, storage and operational work. Confirm the billing currency, taxes and any exchange-rate indexing before making a commitment.
Bring a representative prompt set, target region, expected requests per minute, peak concurrency and context/output limits. Identify sensitive data and retention requirements. Use the model explorer and pricing guide to define the scope before agreeing on capacity and commercial terms.
Latency and throughput depend on the model, input and load. Test them on your workload; availability targets and remedies are defined in the service agreement.
Every GPU runs in Türkiye today. Other territories are served over the network, and new sites are planned.
RTX PRO 6000 Blackwell clusters with 1.3 TB+ of GPU memory. KVKK-oriented data residency and TRY billing.
Served over the network from Türkiye. All GPUs remain in Türkiye.
Compute is hosted in Türkiye today. Frankfurt and Canada are planned locations, not currently available deployment regions.
Current compute is in Türkiye. Frankfurt and Canada are planned; requests from other markets can be served from Türkiye without moving the GPUs.
A shorter network route can reduce transport delay, but first-token time also depends on the model, queue and input size. Compare the same model and prompt from your actual users' location.
No. You still need to assess processing grounds, notices, retention, access and any transfers through integrations or support. Regional compute is one part of that assessment.
Not if the agreed price is indexed to USD or another currency. Confirm whether rates are fixed in TRY, which conversion date applies and whether taxes are included.
Bring a representative prompt set, target region, expected requests per minute, peak concurrency and context/output limits. Identify sensitive data and retention requirements. Use the model explorer and pricing guide to define the scope before agreeing on capacity and commercial terms.
Istanbul, Türkiye · OpenAI compatible · KVKK-oriented