05/Models

Load the model.
Not the infrastructure.

Open-weight, embedding, reranking and speech models behind one API. Bring your own weights when the catalog is not enough.

MetaLlama
QwenQwen
DeepSeekDeepSeek
MistralMistral

Model-family marks identify their respective owners. Availability depends on your deployment; no partnership or endorsement is implied.

ModelFamilyTaskContextDeploymentRegionAvailability
Llama
Chat
128K
Shared · Performance · Dedicated
TR-IST
LIVE
Deploy → API →
Llama
Chat
128K
Shared · Performance
TR-IST
LIVE
Deploy → API →
Llama
Safety · Moderation
128K
Shared · Performance
TR-IST
LIVE
Deploy → API →
Qwen
Chat · Turkish
128K
Shared · Performance · Dedicated
TR-IST
LIVE
Deploy → API →
Qwen
Coding
128K
Shared · Performance · Dedicated
TR-IST
LIVE
Deploy → API →
Qwen
Chat
128K
Shared · Performance
TR-IST
LIVE
Deploy → API →
DeepSeek
Reasoning
128K
Shared · Performance · Dedicated
TR-IST
LIVE
Deploy → API →
DeepSeek
Reasoning
128K
Shared · Performance · Dedicated
TR-IST
LIVE
Deploy → API →
Mistral
Chat
128K
Performance · Dedicated
TR-IST
LIVE
Deploy → API →
Mistral
Chat
32K
Shared · Performance
TR-IST
LIVE
Deploy → API →
Mistral
Chat
32K
Shared
TR-IST
LIVE
Deploy → API →
Turkish
Chat · Turkish
Per model
Shared · Performance · Dedicated
TR-IST
LIVE
Deploy → API →
Utility
Embedding · Turkish
Per model
Shared · Performance
TR-IST
LIVE
Deploy → API →
Utility
Reranking
Per model
Shared · Performance
TR-IST
LIVE
Deploy → API →
Utility
Speech · Turkish
Audio
Shared · Performance
TR-IST
LIVE
Deploy → API →
Custom
Custom
Per model
Dedicated · Enterprise
Your choice
ON REQUEST
Deploy → API →
16 of 16 models 50+ models available · families listed as published at llmnow.ai · availability varies by plan
06/Metal

Under every Pod:
real compute.

LLMPods is not a reseller of someone else's API. Every Pod is a physical stack, from runtime to chassis, and each layer is inspectable.

Compute Pod · exploded view
Pod Runtime stack
RTX PRO 6000 Blackwell
Accelerator · Spec sheet 01
ArchitectureBlackwell
VRAM96 GB GDDR7
DeploymentShared · Performance · Dedicated
RegionTR-IST-01
ReservationReserved capacity
AvailabilityLive · 1.3 TB+ pooled GPU memory
Blackwell · dedicated
Accelerator · Spec sheet 02
ArchitectureBlackwell
VRAMPer class
DeploymentGPU rental · fine-tuning · custom serving
RegionTR-IST-01
Reservation$0.79 – $2.50 / GPU-hour
AvailabilityOn request
Ampere · dedicated
Accelerator · Spec sheet 03
ArchitectureAmpere
VRAMPer class
DeploymentGPU rental · fine-tuning · custom serving
RegionTR-IST-01
Reservation$0.79 – $2.50 / GPU-hour
AvailabilityOn request
Enterprise cluster
Accelerator · Spec sheet 04
ArchitectureCustom
VRAMCustom
DeploymentDedicated cluster · private networking
RegionYour choice
ReservationAnnual contract · 99.9% SLA
AvailabilityCustom agreement
Türkiye cluster: RTX PRO 6000 Blackwell, 1.3 TB+ GPU memory. Rental classes and ranges as published at llmnow.ai. Request Capacity →
07/Regions

Distance is
infrastructure.

Türkiye is physical infrastructure. Other regions are served over the network and are labeled that way: live, served or planned.

RegionStateData residencyBillingNetwork

Illustrative regional structure · confirm states, residency and billing before launch

Regional networkTR-IST-01 · live
TR-IST-01 EUROPE MENA [PLANNED]
Live · hosted Served · routed Planned

Physical infrastructure in Türkiye: RTX PRO 6000 Blackwell clusters with 1.3 TB+ GPU memory. KVKK-oriented data residency, TRY billing and e-Fatura support apply.

Türkiye today. More regions planned.

Compute is hosted in Türkiye today. Frankfurt and Canada are planned locations, not currently available deployment regions.

  1. Live

    Istanbul / Türkiye

    Current compute location · TR-IST-01

  2. Planned

    Frankfurt / Germany

    Planned · no launch date announced

  3. Planned

    Canada

    Planned · no launch date announced

08/Control

Your workload.
Your boundary.

A Dedicated Pod can be cut off from every route except the one you own.

  • Regional processingChoose where requests are processed.
  • Zero-retention optionsRetention settings per deployment.
  • Dedicated infrastructureGPU capacity not shared with other tenants.
  • Private endpointsReach your Pod without the public path.
Public internet
Shared pool
Other tenants
Your network · private endpoint

Boundary active · external routes closed

  • Customer data controlsYou decide what is logged and kept.
  • No training on customer dataYour prompts and outputs stay yours.
  • Custom enterprise agreementsTerms written around your requirements.
  • KVKK-oriented infrastructureDeployment options aligned to KVKK obligations.

Local hosting alone does not establish KVKK compliance.

●/Request trace · live demo

Where does my request go?

Send a request and watch it cross the stack. Every stage reports a timestamp, a duration and a state.

REQUEST7F39A21
FIRST TOKENDEMO 24ms
TOTALDEMO 932ms

Demonstration data

StageDetailTimestampDurationState
Gateway api.llmpods.com · auth ok T+000ms 6ms DONE
Region TR-IST-01 selected T+006ms 5ms DONE
Router route assigned · auto T+011ms 3ms DONE
Model qwen-coder · loaded T+014ms 4ms DONE
Pod POD-018 ready T+018ms 3ms DONE
GPU PRO 6000 · allocation held T+021ms 3ms DONE
Generation first token · streaming T+024ms 908ms DONE
Response stream closed · 200 OK T+932ms 1ms DONE