logo
ホーム 事例

DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories

認証
中国 Beijing Qianxing Jietong Technology Co., Ltd. 認証
中国 Beijing Qianxing Jietong Technology Co., Ltd. 認証
顧客の検討
北京Qianxing Jietongの技術Co.、株式会社の販売スタッフは非常に専門および忍耐強い。それらは引用語句をすぐに提供してもいい。プロダクトの質そして包装はまた非常によい。私達の協同は非常に滑らかである。

—— 《のFestfing DVの》 LLC

私がIntel CPUおよび東芝SSDを緊急に捜していたときに、北京Qianxing Jietongの技術Co.、株式会社からのサンディは私に多くの助けを与え、私に私がすぐに必要としたプロダクトを得た。私は実際に彼女を認める。

—— キティ円

北京Qianxing Jietongの技術Co.、株式会社のサンディは私がサーバーを買う時間の構成間違いを私に思い出させることができる非常に注意深いセールスマンである。エンジニアはまた非常に専門で、すぐにテスト プロセスを完了できる。

—— Strelkin Mikhail Vladimirovich

北京千星捷通との仕事は大変満足しています。製品の品質は素晴らしく、納期も常に守られています。営業チームはプロフェッショナルで、忍耐強く、私たちの質問にすべて丁寧に対応してくれます。彼らのサポートに心から感謝しており、長期的なパートナーシップを期待しています。強くお勧めします!

—— アフマド・ナビド

品質: 提供者との素晴らしい経験. MikroTik RB3011は既に使用されていましたが,非常に良い状態で,すべてが完璧に動作しています. コミュニケーションは迅速でスムーズでした.そして私の懸念はすぐに解決されました信頼性の高いサプライヤーです 強くお勧めします

—— ゲラン・コレシオ

オンラインです

DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories

July 17, 2026
At Paris’s RAISE Summit, DDN showcased its ongoing collaboration with Nebul, a European sovereign hybrid cloud provider, focused on boosting efficiency for large-scale AI inference deployments. Unveiled last week, the joint initiative unites Nebul’s inference platform, DDN’s Infinia data intelligence architecture, and NVIDIA accelerated computing to resolve a key production AI bottleneck: data movement costs and performance limitations during inference workloads.

最新の会社の事例について DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories  0

DDN frames the collaboration around critical production metrics: GPU utilization, token throughput, cost per token, and latency. While model training builds AI asset value, inference defines its operational and commercial returns. The rising adoption of agentic AI, retrieval-augmented generation (RAG), and high-concurrency inference means storage and data infrastructure directly impact accelerator efficiency and AI response speeds.

This active proof-of-concept project has yielded promising early results. The partners have recorded measurable improvements in time-to-first-token with KV cache enabled and completed validation for RoCE-based infrastructure. Ongoing benchmarking covers longer inference sequence lengths, unlocking further optimization potential for the Infinia platform. The collaboration also expands to joint NVIDIA efforts on benchmarking frameworks, scalability verification, and upcoming technical publications.

The integrated platform leverages distributed KV cache services, GPU-native data movement, intelligent data orchestration, and high-performance storage architecture. KV cache acceleration delivers notable inference gains by preserving and rapidly retrieving pre-computed attention states, cutting redundant calculations and eliminating data delivery delays that cause GPU idling.

最新の会社の事例について DDN and Nebul Validate KV Cache Acceleration for NVIDIA-Based AI Factories  1

Leaders from DDN, Nebul, and NVIDIA highlighted a major industry shift: AI infrastructure priorities are moving from raw GPU deployment to operational efficiency, maximizing returns from existing accelerator hardware. DDN CEO Alex Bouzari and Nebul CEO Arnold Juffer noted that past focus on larger model scales has given way to optimizing inference economics to make production AI commercially viable via lower per-token costs. NVIDIA Cloud Infrastructure VP Rod Evans added that large-scale agentic workloads now measure infrastructure success by GPU utilization and latency, rather than sheer compute power.

DDN emphasizes that AI infrastructure must evolve beyond basic storage functions to actively support AI execution workflows. The firm’s infrastructure platforms currently power over one million GPUs worldwide, serving hyperscalers, cloud providers, enterprises, governments, and research institutions.

Modern AI infrastructure teams now prioritize these core production metrics:
GPU utilization: Measures effective accelerator activity during inference, maximizing value of high-end GPU hardware.
Cost per token: Links infrastructure performance directly to AI model output operational costs.
Tokens per watt: Evaluates energy efficiency of AI inference output.
Time to first token: Determines interactive AI application responsiveness and user experience.
Time to production: Quantifies operational effort to migrate AI services from testing to scalable commercial deployment.

As inference becomes the dominant AI workload, delivering cached context and enterprise data to GPUs with low, stable latency will be critical to sustaining high GPU utilization and controlling long-term operational costs.

Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!

連絡先の詳細
Beijing Qianxing Jietong Technology Co., Ltd.

コンタクトパーソン: Ms. Sandy Yang

電話番号: 13426366826

私達に直接お問い合わせを送信 (0 / 3000)