05/Models

Open models
for LLM inference.

Open-weight, embedding, reranking and speech models behind one API. Bring your own weights when the catalog is not enough.

MetaLlama
QwenQwen
DeepSeekDeepSeek
MistralMistral

Model-family marks identify their respective owners. Availability depends on your deployment; no partnership or endorsement is implied.

ModelFamilyTaskContextDeploymentRegionAvailability
Llama
Chat
128K
Shared · Performance · Dedicated
TR-IST
ON REQUEST
Deploy → API →
Llama
Chat
128K
Shared · Performance
TR-IST
ON REQUEST
Deploy → API →
Llama
Safety · Moderation
128K
Shared · Performance
TR-IST
ON REQUEST
Deploy → API →
Qwen
Chat · Turkish
128K
Shared · Performance · Dedicated
TR-IST
ON REQUEST
Deploy → API →
Qwen
Coding
128K
Shared · Performance · Dedicated
TR-IST
ON REQUEST
Deploy → API →
Qwen
Chat
128K
Shared · Performance
TR-IST
ON REQUEST
Deploy → API →
DeepSeek
Reasoning
128K
Shared · Performance · Dedicated
TR-IST
ON REQUEST
Deploy → API →
DeepSeek
Reasoning
128K
Shared · Performance · Dedicated
TR-IST
ON REQUEST
Deploy → API →
Mistral
Chat
128K
Performance · Dedicated
TR-IST
ON REQUEST
Deploy → API →
Mistral
Chat
32K
Shared · Performance
TR-IST
ON REQUEST
Deploy → API →
Mistral
Chat
32K
Shared
TR-IST
ON REQUEST
Deploy → API →
Turkish
Chat · Turkish
Per model
Shared · Performance · Dedicated
TR-IST
ON REQUEST
Deploy → API →
Utility
Embedding · Turkish
Per model
Shared · Performance
TR-IST
ON REQUEST
Deploy → API →
Utility
Reranking
Per model
Shared · Performance
TR-IST
ON REQUEST
Deploy → API →
Utility
Speech · Turkish
Audio
Shared · Performance
TR-IST
ON REQUEST
Deploy → API →
Custom
Custom
Per model
Dedicated · Enterprise
Your choice
ON REQUEST
Deploy → API →
16 of 16 models Illustrative model catalog. Confirm the deployed version, API identifier, context limit and availability before launch.

How to choose an open model

Start with the task and your evaluation data, then select a model and deployment. Parameter count, context length and hosting location describe different constraints; none is a quality score on its own.

Task and language quality

Test representative Turkish and English prompts, domain terminology, citations and structured outputs. Use a fixed evaluation set to compare answers, coding results or extraction accuracy. A family name alone does not determine suitability.

Context and GPU memory

Context includes input and generated tokens. Long prompts and concurrent requests increase KV-cache memory, alongside the model weights. Confirm the deployed context limit, precision and concurrency rather than treating a model-card maximum as guaranteed capacity.

Licence and deployment fit

Open weights do not mean identical commercial rights. Review the specific model licence, acceptable-use terms and fine-tune permissions. Confirm API identifiers, tool support and deployment availability before integrating production traffic.

06/Metal

Under every Pod:
real compute.

LLMPods is not a reseller of someone else's API. Every Pod is a physical stack, from runtime to chassis, and each layer is inspectable.

Compute Pod · exploded view
Pod Runtime stack
RTX PRO 6000 Blackwell
Accelerator · Spec sheet 01
ArchitectureBlackwell
VRAM96 GB GDDR7
DeploymentShared · Performance · Dedicated
RegionTR-IST-01
ReservationReserved capacity
AvailabilityLive · 1.3 TB+ pooled GPU memory
Blackwell · dedicated
Accelerator · Spec sheet 02
ArchitectureBlackwell
VRAMPer class
DeploymentGPU rental · fine-tuning · custom serving
RegionTR-IST-01
Reservation$0.79 – $2.50 / GPU-hour
AvailabilityOn request
Ampere · dedicated
Accelerator · Spec sheet 03
ArchitectureAmpere
VRAMPer class
DeploymentGPU rental · fine-tuning · custom serving
RegionTR-IST-01
Reservation$0.79 – $2.50 / GPU-hour
AvailabilityOn request
Enterprise cluster
Accelerator · Spec sheet 04
ArchitectureCustom
VRAMCustom
DeploymentDedicated cluster · private networking
RegionYour choice
ReservationSLA by agreement
AvailabilityCustom agreement
Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement. Request Capacity →
07/Regions

Distance is
infrastructure.

Türkiye is physical infrastructure. Other regions are served over the network and are labeled that way: live, served or planned.

RegionStateData residencyBillingNetwork

Illustrative regional structure · confirm states, residency and billing before launch

Regional networkTR-IST-01 · live
TR-IST-01 EUROPE MENA [PLANNED]
Live · hosted Served · routed Planned

Physical infrastructure in Türkiye: RTX PRO 6000 Blackwell clusters with 1.3 TB+ GPU memory. KVKK-oriented data residency, TRY billing and e-Fatura support apply.

Türkiye today. More regions planned.

Compute is hosted in Türkiye today. Frankfurt and Canada are planned locations, not currently available deployment regions.

  1. Live

    Istanbul / Türkiye

    Current compute location · TR-IST-01

  2. Planned

    Frankfurt / Germany

    Planned · no launch date announced

  3. Planned

    Canada

    Planned · no launch date announced

08/Control

Your workload.
Your boundary.

A Dedicated Pod can be cut off from every route except the one you own.

  • Regional processingChoose where requests are processed.
  • Zero-retention optionsRetention settings per deployment.
  • Dedicated infrastructureGPU capacity not shared with other tenants.
  • Private endpointsReach your Pod without the public path.
Public internet
Shared pool
Other tenants
Your network · private endpoint

Boundary active · external routes closed

  • Customer data controlsYou decide what is logged and kept.
  • Define permitted data use in your agreementYour prompts and outputs stay yours.
  • Custom enterprise agreementsTerms written around your requirements.
  • KVKK-oriented infrastructureDeployment options aligned to KVKK obligations.

Local hosting alone does not establish KVKK compliance.

●/Request trace · live demo

Where does my request go?

Send a request and watch it cross the stack. Every stage reports a timestamp, a duration and a state.

REQUEST7F39A21
FIRST TOKENDEMO 24ms
TOTALDEMO 932ms

Demonstration data

StageDetailTimestampDurationState
Gateway api.llmpods.com · auth ok T+000ms 6ms DONE
Region TR-IST-01 selected T+006ms 5ms DONE
Router route assigned · auto T+011ms 3ms DONE
Model qwen-coder · loaded T+014ms 4ms DONE
Pod POD-018 ready T+018ms 3ms DONE
GPU PRO 6000 · allocation held T+021ms 3ms DONE
Generation first token · streaming T+024ms 908ms DONE
Response stream closed · 200 OK T+932ms 1ms DONE