AI Pulse by Inblix

SenseTime Claims 2.42T Daily Tokens, But Its New AI Chip Push Is Unverified

AI News · Jul 22, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: SenseTime Claims 2.42T Daily Tokens, But Its New AI Chip Push Is Unverified

SenseTime is making a big bet on domestic chips with its Galaxy Project, a broad coalition of nearly 20 partners aimed at building an AI infrastructure loop from silicon to commercial deployment. Yang Fan, the company’s co-founder and large device group president, announced the initiative alongside a space computing deal with Guoxing Aerospace and a research pact with five institutions including the Shanghai AI Lab. The pitch hinges on a moment where enterprise token demand is exploding and Chinese-made chips are finally good enough to build intelligent computing centers around quickly.

But the numbers SenseTime is selling require squinting. The company says its platform processes 2.42 trillion tokens daily and projects that will skyrocket 25-fold to 10 trillion per day by Q4 2026. That’s a forecast, not an audit. Cost-effectiveness data is similarly self-reported. SenseTime claims its hybrid inference technology boosts model FLOPs utilization by 85–152% on domestic chips and delivers 1.25x the cost-effectiveness of Nvidia’s H-series parts. They also say optimized clusters can double token output at the same cost compared to standard domestic setups, theoretically clearing a profitability hurdle that has kept Chinese silicon in second place.

Anyone who’s operated production AI infrastructure knows a vendor’s sanitized test cluster is a different beast from a customer’s messy real-world pipeline. The multi-chip software fragmentation that has plagued domestic alternatives gets addressed here with a full-stack adaptation layer meant to let workloads move between vendors like Cambricon, Hygon, and Huawei Ascend without a rewrite. SenseTime points to a 3x speedup in a protein prediction task and a 93% multi-card parallel acceleration rate for video generation as proof points. These demos are table stakes. The real verdict won’t come until a paying customer runs a mixed-hardware generation on a live service.

Then there’s the new metric: Tokens Per Watt. It’s a branding exercise on a real problem — data center efficiency — packaged with an orchestration agent that juggles electricity pricing, load forecasting, and scheduling. SenseTime claims an 80% jump in token output per electricity cost, power prices 10% below regional averages, and 96% load prediction accuracy. Those figures look clean in a pilot. I’ll wait for them to survive a summer heatwave and volatile energy markets before calling this a solved problem.

💡 Key Takeaways

  1. Every performance and efficiency figure SenseTime is promoting — from token throughput to cost-effectiveness over Nvidia H-series — is self-reported and lacks third-party validation.
  2. The claim of a 25x token growth by Q4 2026 is a corporate forecast, not a measured trend, and should be treated as aspirational until proven with quarterly results.
  3. The Galaxy Project’s real significance isn't the benchmarks; it’s the roster of domestic chip partners (Cambricon, Hygon, Biren) signaling a collective push to build a viable non-Nvidia ecosystem.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles