09/Workloads
What are you building?
Pick a workload and see the Pod stack it usually needs, which deployment fits, and which model categories sit inside it.
Generated architecture · RAG
EXTERNALUser
APILLMPods API
PODEmbedding Pod
EXTERNALVector database
PODReranker Pod
PODLLM Pod
EXTERNALResponse
10/Observability
Every request
leaves a trace.
Latency, tokens, GPU load and cost, sliced by region, model and Pod. Illustrative figures are marked as demonstration data.
Trace · 7F39A21Demo
00msRequest accepted
06msRegion selected
11msRoute assigned
18msPod ready
24msFirst token
Streaming · 908ms
932msComplete
Requests1.28M
Tokens214M
TTFT24ms
P50 latency0.93s
P95 latency1.9s
GPU utilization67%
Error rate0.02%
Cost₺ [—]
Slice by Region · Model · Pod
Demonstration data
●/Control plane
The same system,
as software.
The console speaks the website's language: the same Pods, status dots, region codes and telemetry. Open the full console on its own artboard.
Active Pods8of 16
Requests / min4,312demo
Tokens / sec182Kdemo
GPU capacity67%
Latency P500.93sdemo
Monthly usage214Mtokens
Requests / min24h
Region health
TR-IST-01Operational
EuropeServed
MENAServed
Pods
PodModelGPUStateLoad
POD-018Llama 70BPRO 6000Active
POD-024Qwen CoderPRO 6000Warm
POD-031CustomPRO 6000Dedicated
POD-041EmbeddingPRO 6000Standby
API keys+ New key
Production
llmp_live_••••••••••3F82
InferenceEmbeddingsModelsUsage
EXPIRATIONNo expiry
CREATEDOct 05
Conceptual UI · demonstration data
Open the full console →
11/Economics
See compute
as it happens.
Inference billing should be as inspectable as the request itself. Tokens in, tokens out, GPU time, and the meter that follows.
Live meter
Prompt
Explain what a Pod is in one paragraph.
Response · streaming
A Pod is an isolated compute unit that hosts a model on reserved GPU capacity. Requests are routed to a Pod, scheduled onto the GPU, and tokens stream back as they are generated. Under load Pods replicate; when traffic falls they return to standby.
REGIONTR-IST
MODELLLAMA-70B
PODPOD-018
INPUT TOKENS412
OUTPUT TOKENS1,284
GPU TIME919 ms
Estimated cost
(412 + 1,284) tokens × $0.25 / 1M = $0.000424
Illustrative meter · Starter rate, 8K context · rates as published at llmnow.ai
Where applicable in Türkiye
- TRY billingInvoiced in Turkish lira.
- e-Fatura / e-Arşiv supportInvoices that fit local accounting workflows.
- Local commercial relationshipContracts and support with a Türkiye-based team.
- Usage analyticsPer key, per model, per Pod, exportable.