InferenceX Puts GB300 First for MiniMax M2.7, but Only on a Fixed Chat Test
The new ranking gives infrastructure teams a comparable speed and cost snapshot, but its single-turn design does not test the multi-step agent work MiniMax highlights for M2.7.
Listen to this story
The audio brief
Story brief
3 key pointsInferenceX’s latest MiniMax M2.7 ranking puts NVIDIA’s GB300 NVL72 at 13,707 tokens per second per GPU, 10% ahead of GB200, but only under a tightly matched serving test: one 8,000-token prompt, 1,000-token response, and 50-token-per-second-per-user target. FP4 and different serving stacks are part of the result, so it is not a general agent-performance verdict. GB200 is cheaper on listed token cost ($0.041 versus...
- 01
GB300 led B300, B200 and MI355X, which recorded 12,253, 11,606 and 7,692 tokens per second per GPU.
- 02
InferenceX excluded platforms without measurements at the 50-token-per-second-per-user operating point.
- 03
GB300 and GB200 used Dynamo vLLM; B300, B200 and MI355X used vLLM, all at FP4.
MiniMax pitches M2.7 for complex agent harnesses and software-engineering work. InferenceX’s new ranking answers a more contained question: on a single-turn chat workload, GB300 NVL72 delivered the highest measured throughput at 13,707 tokens per second per GPU.
GB200 NVL72 came second at 12,515 tokens per second per GPU, a 10% gap. B300, B200 and MI355X followed at 12,253, 11,606 and 7,692 tokens per second per GPU, respectively.
A shared responsiveness constraint
Each platform was read at a matched target of 50 tokens per second per user. The test used one chat turn with 8,000 input tokens and 1,000 output tokens. InferenceX excludes hardware without a measurement at that operating point, rather than comparing a system at a different user-speed tradeoff.
The result includes the serving stack
This is not a hardware-only comparison. All listed measurements use FP4 precision, while GB300 and GB200 use Dynamo vLLM; B300, B200 and MI355X use vLLM. InferenceX says it measures real hardware, sweeps concurrency to trace each platform’s throughput-versus-interactivity frontier, and reruns tests as engines and configurations change.
A speed result, not an agent verdict
MiniMax announced M2.7 in March and describes it as supporting complex agent harnesses, agent teams, dynamic tool search and software-engineering tasks. Those are broader patterns than the ranking’s one prompt-and-response exchange. The benchmark therefore establishes a leader for its specified serving case, not across every workload MiniMax associates with M2.7.
Speed and listed token cost also point in different directions. InferenceX puts GB300 at $0.047 per million tokens, versus $0.041 for GB200, using measured throughput and GPU-hour rates from SemiAnalysis’s AI Cloud total-cost-of-ownership model. The page reruns continuously, but its newest input result was dated June 13, so later engine releases or configurations could change the order.
Sources
- inferencex.semianalysis.comFastest GPU for MiniMax M2.7 Inference: Live Rankings | InferenceX
- minimax.ioMiniMax M2.7: Early Echoes of Self-Evolution