09/Workloads

LLM workloads
from RAG to agents.

Pick a workload and see the Pod stack it usually needs, which deployment fits, and which model categories sit inside it.

Generated architecture · RAG
EXTERNALUser
APILLMPods API
PODEmbedding Pod
EXTERNALVector database
PODReranker Pod
PODLLM Pod
EXTERNALResponse

What each workload actually needs

These architecture guides describe application responsibilities, model roles and evaluation criteria. They are reference patterns; your data, traffic and integration requirements determine the final configuration.

AI assistants

A chat model handles dialogue while your application manages session history, authentication and tool permissions. Limit retained history and define escalation paths. Evaluate task completion, first-token latency and behaviour when information is missing.

Customer support

Connect an approved knowledge base and pass only the customer information needed for the task. Ticket creation and account changes require application-side authorization. Measure correct answers, human handoffs and policy adherence, not just conversation volume.

RAG and document search

Your application chunks documents and maintains retrieval permissions and a vector database. Embeddings retrieve candidates; optional reranking improves selection before the LLM produces an answer. Evaluate source citations, retrieval recall and unsupported claims.

Document intelligence

Extract text with OCR or a suitable document parser before requesting structured fields from an LLM. Validate output against a schema and retain source references. Review critical contractual or financial fields with a qualified person.

Coding assistants

Select a code-capable model and control which repository files enter the prompt. Keep secrets out of context and execute generated code in a controlled environment. Compare test pass rates, edit quality and response time on your own codebase.

Classification and extraction

Define labels and a JSON schema before processing a batch. Use representative examples and an explicit unknown category. Track precision, recall, schema-valid outputs and cost per record; route low-confidence results for review.

Speech and call transcription

A speech-to-text model converts audio; a separate language model can summarize or extract actions. Confirm audio format, consent and retention before ingestion. Evaluate word error rate, Turkish accents and speaker handling on real recordings.

Tool-using agents

The LLM proposes tool calls; your application validates arguments and executes allowed actions. Set permission boundaries, retry limits and audit logs. Evaluate successful completion, unintended actions and recovery from tool failures.

10/Observability

Every request
leaves a trace.

Latency, tokens, GPU load and cost, sliced by region, model and Pod. Illustrative figures are marked as demonstration data.

Trace · 7F39A21Demo
00msRequest accepted
06msRegion selected
11msRoute assigned
18msPod ready
24msFirst token
Streaming · 908ms
932msComplete
Requests / minLast 24h · demo data
6K4.5K3K1.5K0 PEAK 5.2K 00:0006:0012:0018:0024:00
LatencyP50P95
2s1s0 P50 · DEMO 0.93s P95 · DEMO 1.9s 00:0006:0012:0018:0024:00
Requests1.28M
Tokens214M
TTFT24ms
P50 latency0.93s
P95 latency1.9s
GPU utilization67%
Error rate0.02%
Cost₺ [—]
Slice by  Region · Model · Pod Demonstration data
●/Control plane

The same system,
as software.

The console speaks the website's language: the same Pods, status dots, region codes and telemetry. Open the full console on its own artboard.

LLMPods / Production / TR-IST-01 All systems operational
Active Pods8of 16
Requests / min4,312demo
Tokens / sec182Kdemo
GPU capacity67%
Latency P500.93sdemo
Monthly usage214Mtokens
Requests / min24h
Region health
TR-IST-01Operational
EuropeServed
MENAServed
Pods
PodModelGPUStateLoad
POD-018Llama 70BPRO 6000Active
POD-024Qwen CoderPRO 6000Warm
POD-031CustomPRO 6000Dedicated
POD-041EmbeddingPRO 6000Standby
API keys+ New key
Production
llmp_live_••••••••••3F82
InferenceEmbeddingsModelsUsage
EXPIRATIONNo expiry
CREATEDOct 05
Conceptual UI · demonstration data Open the full console →
11/Economics

See compute
as it happens.

Inference billing should be as inspectable as the request itself. Tokens in, tokens out, GPU time, and the meter that follows.

Live meter
Prompt
Explain what a Pod is in one paragraph.
Response · streaming
A Pod is an isolated compute unit that hosts a model on reserved GPU capacity. Requests are routed to a Pod, scheduled onto the GPU, and tokens stream back as they are generated. Under load Pods replicate; when traffic falls they return to standby.
REGIONTR-IST
MODELLLAMA-70B
PODPOD-018
INPUT TOKENS412
OUTPUT TOKENS1,284
GPU TIME919 ms
Estimated cost (412 + 1,284) tokens × $0.25 / 1M = $0.000424

Illustrative rates and limits, not a current offer. Final model availability, price, capacity, retention and SLA are defined in your deployment agreement.

Where applicable in Türkiye
  • TRY billingInvoiced in Turkish lira.
  • e-Fatura / e-Arşiv supportInvoices that fit local accounting workflows.
  • Local commercial relationshipContracts and support with a Türkiye-based team.
  • Usage analyticsPer key, per model, per Pod, exportable.
See how it is packaged →