Skip to content
Platform overview

Performance

Keep AI responsive as the team grows.

A model runtime is one part of the system. Kaldryn adds workload routing, shared inference capacity and the measurements that help administrators understand latency, throughput and hardware pressure.

Explore this part of the platform
01

Route the work

Configure fast, mid, large, vision, CPU and fallback model roles. Inspect routing decisions, escalation and downgrade activity as requests move through the available capacity.

02

Measure the experience

Follow time to first token, tokens per second and context-fill percentiles. Compare benchmark runs with live resource readings and inspect request activity across users and model tiers.

03

Use the hardware deliberately

See GPU memory, CPU, storage and pending work. Manage model placement and GPU nodes through the cluster console, with secure node enrollment, connectivity checks and engine configuration.

Explore this part of the platform

Kaldryn Platform · Performance

Interactive software preview · illustrative data

Performance

Measure inference performance on your deployment.

Illustrative data · safe to explore

TTFT p50

180 ms

Tokens / sec

420

Active requests

32

Live metrics

Time to first token benchmarks

Preview only · no changes to your systems

What this module includes

  • Throughput and latency
  • Resource utilisation
  • Performance diagnostics
  • Time to first token benchmarks
  • Per-tier live serving metrics
  • Context-fill p50 and p95
  • Escalated and downgraded routing decisions

Based on the Kaldryn Platform admin console. Available features depend on license tier, role, hardware and configuration.

Throughput and concurrency vary by model, quantisation, context length, hardware and workload. The comparison below illustrates a conservative 9× scenario; confirm performance with your own workload.

More from the same GPU

More work. Same hardware.

See the potential of shared inference with a conservative 9× comparison for 32 simultaneous chats on GB10 hardware.

0×

the per-user throughput in this illustration

Inside the Kaldryn Engine
Illustrative tokens per second, per user
GB10 · 0 concurrent chats
Kaldryn engine

Continuous batching · shared GPU

0.0tok/s
Standard Ollama

Reference baseline for this illustration

0.0tok/s
Illustration · 32 simultaneous chats · GB10
01

Performance for shared workloads

At 6.3 tokens per second per user versus a 0.7 reference baseline, the illustrated throughput is 9×. Across 32 chats, that is 201.6 versus 22.4 tokens per second.

02

A workspace, ready for your people

Chat, document search with RAG, agents and model fine-tuning come together in one platform.

03

Administration included

Manage SSO, role-based access, DLP and audit logs alongside GPU health, signed updates and backups in one console.

Illustrative comparison, not a new measured benchmark: 0.7 × 9 = 6.3 tokens/s per user. Reference workload: GB10, Qwen2.5-7B, 8k context, 150-token responses, warm engine, 32 concurrent chats. Actual throughput depends on the model, hardware and configuration.

Inside these modules

ModelsChoose, install and operate models on your own infrastructure.Explore capabilities
  • Model library and hardware advisor
  • CPU and GPU deployment
  • Model lifecycle and trust information
  • Search and filter the model catalogue
  • Memory and context requirements
  • Routing defaults by tier
  • Import local model files and GGUF
AnalyticsUnderstand adoption, performance and resource consumption.Explore capabilities
  • Model and user activity
  • Token consumption and latency
  • Usage trends and reporting
  • 7, 30 and 90-day comparisons
  • Daily cloud-cost estimates with rate basis
  • Request distribution by tier
  • Hourly usage heatmap and user breakdown
GPU ClusterOperate compute capacity across your own nodes.Explore capabilities
  • Node registration and health
  • Model placement
  • GPU resource visibility
  • One-time node enrollment tokens
  • Secure Node Link with mTLS
  • GPU count and tensor-parallel sizing
  • Ping, optimise and detach nodes
HardwareSee the resources your AI workloads depend on.Explore capabilities
  • CPU, memory and storage
  • GPU visibility
  • Appliance resource monitoring
  • Per-node GPU selection
  • VRAM used and available
  • Active and pending requests
  • GPU clocks and resource history
PerformanceMeasure inference performance on your deployment.Explore capabilities
  • Throughput and latency
  • Resource utilisation
  • Performance diagnostics
  • Time to first token benchmarks
  • Per-tier live serving metrics
  • Context-fill p50 and p95
  • Escalated and downgraded routing decisions

Discuss your deployment

Discuss your deployment