Home Services Architecture Advantages Scenarios About Member Center
GPU Cluster Online
Models Ready
API Availability 99.99%

200+ Models, Massive Compute, One API for All

DeepSeek·GLM·Qwen·Kimi·Llama Fully Integrated, Enterprise-Grade AI Compute Platform

Model List

All Available Models at a Glance

101+ models pre-integrated, core models listed below, continuously updated

Model Parameters Context Length Precision Input Price (Ref.) Output Price (Ref.) Status
GPT-4.1 mini OpenAI1M¥2.80 /百万Token¥11.20 /百万TokenOnline
GPT-4.1 HOT1M¥14.00 /百万Token¥56.00 /百万TokenOnline
o3 Reasoning200K¥14.00 /百万Token¥56.00 /百万TokenOnline
Claude Sonnet 4.5 Anthropic200K¥21.00 /百万Token¥105.00 /百万TokenOnline
Claude Haiku 3.5 Anthropic200K¥5.60 /百万Token¥28.00 /百万TokenOnline
Gemini 2.5 Pro Google1M¥8.75 /百万Token¥70.00 /百万TokenOnline
Gemini 2.5 Flash Google1M¥1.05 /百万Token¥4.20 /百万TokenOnline
DeepSeek-V4 Pro Latest685B MoE128KBF16¥12.00 /百万Token¥24.00 /百万TokenOnline
DeepSeek-V4 Flash Latest685B MoE128KBF16¥1.00 /百万Token¥2.00 /百万TokenOnline
DeepSeek-R1 Reasoning685B MoE128KBF16¥2.00 /百万Token¥5.00 /百万TokenOnline
GLM-5 智谱200B+128KBF16¥1.00 /百万Token¥6.00 /百万TokenOnline
GLM-5V-Turbo 多模态130B+128KBF16¥1.50 /百万Token¥3.00 /百万TokenOnline
Qwen3.7 Max Latest397B MoE128KBF16¥2.00 /百万Token¥6.00 /百万TokenOnline
Qwen3.6 Flash Fast235B MoE128KBF16¥0.30 /百万Token¥0.60 /百万TokenOnline
Qwen3-Coder-Next Code235B MoE128KBF16¥1.50 /百万Token¥3.50 /百万TokenOnline
Kimi-K2.5 Latest1T MoE256KBF16¥2.50 /百万Token¥5.00 /百万TokenOnline
Kimi-K2 Thinking Reasoning1T MoE256KBF16¥3.00 /百万Token¥6.00 /百万TokenOnline
MiniMax-M2.5 Latest456B MoE1MBF16¥2.00 /百万Token¥4.00 /百万TokenOnline
MiniMax-M2.5 Highspeed Fast456B MoE1MBF16¥1.00 /百万Token¥2.00 /百万TokenOnline
Doubao Seed 2.0 Pro 豆包128KBF16¥1.50 /百万Token¥3.00 /百万TokenOnline
Doubao Seed 2.0 Code Preview Code128KBF16¥1.00 /百万Token¥2.00 /百万TokenOnline
Llama-4 Meta400B MoE1MBF16¥2.50 /百万Token¥5.00 /百万TokenOnline
Yi-Lightning 零一万物64KBF16¥0.50 /百万Token¥1.50 /百万TokenOnline
Step-2 阶跃星辰128KBF16¥1.50 /百万Token¥3.20 /百万TokenOnline
Hy3 Preview 混元128KBF16¥1.00 /百万Token¥2.00 /百万TokenOnline
BAAI/bge-m3 Embedding568M8KFP16¥0.02 /百万TokenOnline
Use Cases

One Command, Deploy Your AI

Explore KHB AI infrastructure deployment solutions across industries, the developer way

enterprise-ai.sh
$ khb deploy --scenario=enterprise --model=DeepSeek-V3.2
算力资源分配完成
模型加载完成 DeepSeek-V3.2
API网关配置完成
企业AI平台已就绪
ENDPOINT: https://api.khb.com/v1/enterprise
STATUS: RUNNING
compute-center.sh
$ khb deploy --scenario=compute --chip=ascend,nvidia
异构算力统一纳管启动
NVIDIA A100 × 128 已接入
昇腾910B × 64 已接入
弹性调度引擎就绪
UTILIZATION: 94.7%
finance.sh
$ khb deploy --scenario=finance --security=private
Private Deployment环境初始化
数据隔离策略已启用
全链路加密传输已开启
Audit Logs系统就绪
COMPLIANCE: PASSED
manufacturing.sh
$ khb deploy --scenario=manufacturing --instance=reserved
预留实例资源锁定
工业AI模型部署完成
7×24 SLA保障已启用
质检模型推理延迟 <50ms
UPTIME: 99.99%
Reserved Instances

Lock Compute, Secure Critical Business

Dedicated Reserved Compute · Model Precision Guaranteed · Cost Controllable · Enterprise SLA

DeepSeek-V3.2
deepseek-ai/DeepSeek-V3.2
Price ¥594,000/组/月
Effective Unit Price ¥2.20/M tokens
TPM 1,250万
TTFT 1,600ms
TPS 45
Context 1M
Suitable for enterprise-grade complex reasoning and decision analysis, code generation and software development assistance, agent tool calling
GLM-5
zai-org/GLM-5
Price ¥594,000/组/月
Effective Unit Price ¥2.75/M tokens
TPM 1,000万
TTFT 1,500ms
TPS 30
Context 1M
Suitable for enterprise-grade agent development, complex task planning and multi-step execution, software engineering automation
Kimi-K2.5
moonshotai/Kimi-K2.5
Price ¥594,000/组/月
Effective Unit Price ¥6.88/M tokens
TPM 400万
TTFT 1,500ms
TPS 30
Context 256K
Suitable for enterprise-grade multimodal agent development, visual content understanding and analysis, complex task automation
MiniMax-M2.5
MiniMaxAI/MiniMax-M2.5
Price ¥297,000/组/月
Effective Unit Price ¥2.75/M tokens
TPM 500万
TTFT 500ms
TPS 30
Context 1M
Suitable for enterprise-grade long document and knowledge base analysis, intelligent customer service and content generation, business process automation
* Effective unit price based on TPM, calculated with 30 days/month and 50% overall utilization. Performance data based on typical inference parameters: 24K input tokens, 1K output tokens, 80% cache hit rate. (Updated May 20, 2026)
AI Gateway

Private LLM Service Gateway

Unified Management · Intelligent Routing · Rate Limiting · Full-Chain Observability

01 Unified Multi-Model Access
One-stop access and standardized calling of different vendor models, supporting OpenAI-compatible interfaces,告别point management, easily managing multi-vendor ecosystems
02 Intelligent Routing & Load Balancing
Intelligent dynamic routing combining traffic characteristics and LLM service characteristics, load balancing, fallback failover, ensuring service stability and business SLA
03 Fine-Grained Governance & Rate Limiting
Configure model permissions, traffic and quota management by user, API Key, project, organization dimensions, multi-tenant isolation, achieving fine-grained model call governance
04 Accurate Cost Accounting
Full-chain cost transparency across consumer users, API Keys, projects, organizations, models and compute, achieving accurate cost accounting and budget control
05 Full-Chain Model Observability
Multi-dimensional observability of model call volume, performance and other metrics, supporting precise model governance, lifecycle management and routing strategy adjustment, supporting A/B testing and canary releases
06 Enterprise Data Security
Bidirectional desensitization real-time privacy risk filtering, linked sensitive content interception and audit logs, ensuring every LLM business transaction is compliant and traceable
Security & Compliance

End-to-End Defense-in-Depth System

Intelligence-driven security, full-chain protection from data to application

🔐
End-to-End Encryption
Full-chain TLS encrypted transmission, data storage encryption, key management system, ensuring absolute security of data in transit and at rest
🛡️
Bidirectional Desensitization
Input/output bidirectional real-time desensitization filtering, automatically identifying and masking sensitive information, preventing privacy data leakage
📋
Audit Logs
Full-chain operation auditing, every API call traceable and auditable, meeting compliance requirements in finance, healthcare and other industries
🏢
Multi-Tenant Isolation
Strict data isolation between tenants, fine-grained permission control, supporting multi-dimensional access control by organization, project and user
🔒
Private Deployment
Data never leaves the domain, all inference completed within the enterprise intranet, meeting the strictest data sovereignty and compliance requirements
🚨
Content Security Detection
实时防御潜在攻击,Content Security Detection准确率超99%,敏感内容自动拦截,保障输出合规
Partnership

Flexible Collaboration Models

Whether you have compute or need compute, KHB provides matching partnership solutions

🤝 Joint Operations
Suitable for those with compute resources who want to quickly gain Token service capabilities and serve end customers together with KHB
Typical Partners
IDC operators, regional intelligent computing centers, GPU cloud service providers, domestic chip manufacturers
Value Benefits
✓ Complete Token production capability, no need to build your own tech team
✓ Under equivalent compute, inference throughput significantly improved
✓ Revenue sharing based on actual service volume
✓ KHB brand endorsement and market support
⚡ Compute Absorption / Compute-as-a-Service
Suitable for those with self-built GPU clusters who want to improve inference efficiency, reduce O&M costs, or convert redundant resources into Token service revenue
Typical Partners
Government/enterprise clients with self-built compute, large internet companies, financial institutions, telecom operators
Value Benefits
✓ Inference efficiency significantly improved, equivalent compute supports larger business scale
✓ GPU performance fully utilized, solving adaptation challenges
✓ Data runs in your own environment, meeting security compliance requirements
✓ Redundant compute can provide Token services externally, generating additional revenue
Technical Architecture

From Compute to Application Data Flow Architecture

Four-layer architecture, data flows left to right, layer-by-layer collaboration, building an enterprise AI infrastructure closed loop

Layer 01
Compute Resource Layer
NVIDIA A100/H100
Huawei Ascend 910B
Muxi GPU
Moore Threads GPU
Layer 02
Inference Service Layer
Model Loading & Scheduling
KV Cache Optimization
Continuous Batching
Quantization & Acceleration
Layer 03
API Gateway Layer
OpenAI-Compatible Interface
Intelligent Routing Gateway
Rate Limiting & Circuit Breaking
Token Metering & Billing
Layer 04
Application Layer
Intelligent Customer Service
Content Generation
Data Analysis
Industry Solutions
0
Pre-integrated Models
0
Chip Architecture Support
0
Inference Latency Reduction %
0
Throughput Boost ×
Core Services

Service Matrix Around KHB·AI Core

Four services orbit the core, from underlying compute to upper gateway, full-chain coverage of enterprise AI infrastructure needs

KHB·AI
INFRASTRUCTURE
Compute Operations
Token Factory
GPU compute operations, efficiently converting GPU resources into Token productivity. Supports NVIDIA and domestic chips, achieving multi-architecture compute unified access and elastic scheduling
🔒
Reserved Instances
Reserved Instances
Dedicated reserved compute, exclusive resources ensuring stable business operation, model precision guaranteed, cost controllable, enterprise SLA commitment
🏗️
MaaS Platform
Enterprise MaaS
One-stop AI solution platform, heterogeneous compute management, model training, inference deployment full-process coverage, 100+ models pre-integrated, ready to use
🌐
AI Gateway
AI Gateway
Private LLM service gateway, unified multi-model access management, intelligent routing and rate limiting, full-chain observability, enterprise data security assured
Compute Operations
TOKEN FACTORY
GPU算力运营,将GPU资源高效转化为Token生产力。支持NVIDIA及国产芯片(昇腾、沐曦、摩尔线程),实现多架构算力统一接入与弹性调度
🔒
Reserved Instances
RESERVED INSTANCES
专属预留算力,独占式资源保障业务稳定运行。模型精度有保障,成本可控,企业级SLA承诺
🏗️
MaaS Platform
ENTERPRISE MAAS
One-stop AI solution platform, heterogeneous compute management, model training, inference deployment full-process coverage, 100+ models pre-integrated, ready to use
🌐
AI Gateway
AI GATEWAY
Private LLM service gateway, unified multi-model access management, intelligent routing and rate limiting, full-chain observability, enterprise data security assured
Advantages

Why Choose KHB

Numbers don't lie, every metric is proof of strength

0
Pre-integrated Models
DeepSeek, GLM, Qwen, Kimi, GPT-4.1, Claude, Gemini and other global mainstream LLMs with one-click access, dynamic model mirror updates, new models adapted immediately, API ready to call
0
Inference Latency Reduction
自研推理加速引擎,KV Cache深度优化与Continuous Batching,首Token延迟降低70%,无损精度动态量化,推理计算量减少60-80%
0
Throughput Boost
异构算力弹性调度,单卡Token产出深度优化,同等算力下Throughput Boost3至5倍,GPU集群利用率提升300%
0
Chip Architecture Support
Simultaneously supports NVIDIA A100/H100, Huawei Ascend 910B, Muxi GPU, Moore Threads GPU, multi-architecture unified management and intelligent scheduling, avoiding single vendor lock-in
0
Service Availability SLA
Enterprise SLA commitment of 99.9% availability, multi-AZ deployment, automatic failover and fallback degradation, 7×24 technical support ensuring business continuity
0
Content Security Detection
Bidirectional Desensitization实时过滤隐私风险,Content Security Detection准确率超99%,联动敏感内容拦截与Audit Logs,数据泄露风险降低99%
0
Quick Start
可视化配置界面,3分钟内完成操作,30+开箱即用预置模板,OpenAI-Compatible Interface,无需深厚技术背景即可跨场景调用
0
Reserved Instance Delivery
Standard reserved instances deployed within 1-7 business days, platform handles model deployment and performance verification, inference performance tuning support, ensuring stable business onboarding
FAQ

FAQ

How strong is KHB's AI infrastructure technical capability?
KHB has self-developed AI compute operations and inference acceleration engines, with core technical capabilities including heterogeneous compute unified management, high-performance model deployment, and intelligent routing scheduling. We have served hundreds of enterprise clients across finance, manufacturing, healthcare and other industries, providing full-chain AI infrastructure services from compute resources to application deployment.
What is the difference between Reserved Instances and pay-as-you-go?
Reserved Instances are dedicated compute resources not shared with others, ensuring model precision and inference stability, suitable for production workloads with strict latency and availability requirements. Pay-as-you-go uses shared resource pools billed by actual usage, suitable for development/testing and elastic business scenarios. Long-term stable use of Reserved Instances is more cost-effective.
Which domestic chips are supported?
Currently supports Huawei Ascend 910B, Muxi GPU, Moore Threads GPU and other domestic chip architectures, as well as NVIDIA A100/H100 and other international mainstream GPUs. Through heterogeneous compute unified management, enterprises can flexibly choose chip solutions, avoiding single vendor lock-in.
How does the AI Gateway ensure data security?
AI网关支持Private Deployment,所有数据在企业内网流转,不经过公网。提供数据脱敏、访问控制、全链路加密传输、操作Audit Logs等企业级安全能力,满足金融、医疗等行业的合规要求。
MaaS平台支持Private Deployment吗?
支持。MaaS平台提供公有云SaaS和Private Deployment两种模式。Private Deployment可将完整平台部署在企业数据中心,数据完全不出域,适合对数据安全有严格要求的行业客户。
How do I get started with KHB's services?
You can contact us through the contact information at the bottom of the page. Our solution team will provide customized proposals based on your business needs. From requirement discussion, solution design, environment deployment to operations and maintenance, we provide professional support throughout the entire process.
Why do enterprises need an AI Gateway if they already have LLM APIs?
As enterprises introduce multiple LLMs, common problems quickly emerge: diverse model sources lead to inconsistent API protocols, scattered call chains lack unified management, different applications have varying SLA requirements that are difficult to meet holistically, and usage volume and costs are hard to measure. The AI Gateway centrally solves these problems, providing unified access, intelligent routing, fine-grained governance and full-chain observability.
Will the AI Gateway add network overhead?
The AI Gateway uses a high-performance proxy architecture, with additional latency typically within 1-3ms, which is negligible compared to the LLM inference latency itself (hundreds of milliseconds to seconds). Meanwhile, the gateway's intelligent routing and connection pool optimization can actually reduce overall latency.
How to control LLM usage costs through the gateway?
The gateway supports configuring traffic quotas and rate limiting policies by user, API Key, project, organization and other dimensions, providing full-chain cost transparency and accurate accounting, helping enterprises monitor LLM usage costs in real-time and avoid budget overruns.
When should an enterprise deploy a private MaaS platform?
Enterprises should consider it when: ① Business involves sensitive data with strict data sovereignty requirements; ② Need to deploy AI capabilities at scale across numerous scenarios with extremely high inference performance and stability requirements; ③ Have diverse domestic or heterogeneous compute environments requiring unified management and efficient utilization; ④ Want to keep up with AI technology development but lack the engineering team for continuous model adaptation and optimization.
How are KHB's international AI and domestic AI services divided and applied?
I. Service Entity Division
KHB INC (US Corporation): Provides full-category overseas AI model API services, exclusively for overseas enterprises and individuals with overseas identity. Access requirements: complete qualification certification, supports USD credit card, USD, and USDT settlement. This service is limited to overseas scenarios and must not be improperly transferred domestically.
Hangzhou Kaihabesi Ecological Technology Co., Ltd.: Primarily operates domestic compliant LLM proxy API services, with all data processed domestically, strictly adhering to all domestic laws and regulations.

II. Rights, Responsibilities & Compliance Declaration
The two entities operate independently, each bearing legal responsibility for their respective services, with no joint liability.
Overseas AI service-related activities follow local regulations. Users are strictly prohibited from using overseas interfaces to deliver services domestically or improperly transmitting domestic data. Violators bear the consequences themselves.
Domestic services are strictly prohibited from being used in illegal or non-compliant scenarios. Users must comply with national regulatory requirements.
The platform reserves the right to verify usage qualifications and may suspend or terminate service for non-compliant accounts.

KHB INC
Hangzhou Kaihabesi Ecological Technology Co., Ltd.
Launch Your AI Infrastructure

Whether you're exploring AI for the first time or seeking compute and model service upgrades, KHB provides professional solutions

Launch LAUNCH
Business Inquiry: ai@khb.hk