Performance for shared workloads
At 6.3 tokens per second per user versus a 0.7 reference baseline, the illustrated throughput is 9×. Across 32 chats, that is 201.6 versus 22.4 tokens per second.
Local & Self-hosted AI for secure chat, agents and knowledge with model governance, DLP and audit evidence for enterprises.
Built for the teams responsible for trust
A workspace for your people. A control plane for your IT team. An evidence layer for security and compliance. All on your infrastructure.
Deploy in your private environment. Choose models and configure the platform around your workflows.
Bring in your identity provider, document libraries and internal systems. Define who can use each model and tool.
Manage policies, review activity, verify network isolation and collect evidence from one administration console.
Interactive software preview · sample data · explore all 30 modules
An operational overview of your AI estate.
Requests
1,284
GPU nodes
3
Ready services
12/12
GPU node health
Based on the Kaldryn Platform admin console. Available features depend on license tier, role, hardware and configuration.
Local hosting is the starting point. Kaldryn adds the identity, governance and operational tools needed to make private AI accountable.
Privacy becomes a workflow.
Limit exposure with local processing, scoped document access and configurable DLP. Review and execute subject-data erasure requests with an audit trail.
Make responsible AI operational.
Make AI use transparent with in-product disclosures, model governance and operational oversight. Keep policies, audit activity and reviewable evidence together.
Bring SAML SSO, role-based access, MFA and passkeys together with network-isolation checks, an egress register and reviewable security evidence.
Kaldryn provides technical controls and compliance-supporting documentation, not automatic regulatory compliance or certification. Your obligations depend on the use case, model, configuration and organisational measures.
Discuss your security requirementsCloud subscriptions and token charges add up across a team. With Kaldryn, local inference has no token or credit bill. Compare your cloud spend with the full cost of running your own AI.
Cost comparison · USD
Include subscriptions and average token or credit spend.
Total Kaldryn cost / month
Fixed at $7,500 per month, with unlimited users to onboard.
Potential annual savings
No metered token or credit charges for local inference.
This comparison uses a fixed Kaldryn budget of $7,500/month. Adjust the team size and cloud spend to reflect your organisation. Cloud defaults are illustrative. Model quality and deployment capacity vary; confirm the configuration and scope in your quote.
Cloud = people × monthly cost × 12. Kaldryn = total monthly cost × 12.
See the potential of shared inference with a conservative 9× comparison for 32 simultaneous chats on GB10 hardware.
the per-user throughput in this illustration
Continuous batching · shared GPU
Reference baseline for this illustration
At 6.3 tokens per second per user versus a 0.7 reference baseline, the illustrated throughput is 9×. Across 32 chats, that is 201.6 versus 22.4 tokens per second.
Chat, document search with RAG, agents and model fine-tuning come together in one platform.
Manage SSO, role-based access, DLP and audit logs alongside GPU health, signed updates and backups in one console.
Illustrative comparison, not a new measured benchmark: 0.7 × 9 = 6.3 tokens/s per user. Reference workload: GB10, Qwen2.5-7B, 8k context, 150-token responses, warm engine, 32 concurrent chats. Actual throughput depends on the model, hardware and configuration.
Tell us about your company, your users and how you use AI today. We’ll help you compare costs and find the right setup.
Company name and business email required.