InferenceX Puts GB300 First for MiniMax M2.7, but Only on a Fixed Chat Test

The new ranking gives infrastructure teams a comparable speed and cost snapshot, but its single-turn design does not test the multi-step agent work MiniMax highlights for M2.7.

By 2 min read
InferenceX Puts GB300 First for MiniMax M2.7, but Only on a Fixed Chat Test
InferenceX Puts GB300 First for MiniMax M2.7, but Only on a Fixed Chat Test

Listen to this story

The audio brief

About 1:40
0:001:40
Read transcript
GB300 NVL72 has taken the top spot in InferenceX’s latest MiniMax M2.7 ranking, delivering 13,707 tokens per second per GPU. That is roughly 10 percent faster than GB200 NVL72, at 12,515 tokens per second. But the important qualifier is the test. InferenceX measured one chat exchange: an 8,000-token prompt followed by a 1,000-token response, while requiring a responsiveness target of 50 tokens per second per user. Platforms without a measurement at that exact operating point were left out, so this is a controlled snapshot—not a universal speed leaderboard. The result also reflects the software stack. Both GB300 and GB200 ran Dynamo vLLM, while B300, B200, and AMD’s MI355X used vLLM. Every system ran at FP4 precision. B300 reached 12,253 tokens per second per GPU, B200 reached 11,606, and MI355X reached 7,692. That makes GB300 the throughput leader for this specific serving case. It does not establish a winner for the broader workloads MiniMax associates with M2.7, including agent teams, dynamic tool search, and software engineering, because those multi-step patterns were not tested. Cost points the other way: InferenceX lists GB300 at 4.7 cents per million tokens, versus 4.1 cents for GB200. The ranking is rerun as engines and configurations change, and its newest input measurement was dated June 13—so the open question is how long this ordering holds.

Story brief

3 key points

InferenceX’s latest MiniMax M2.7 ranking puts NVIDIA’s GB300 NVL72 at 13,707 tokens per second per GPU, 10% ahead of GB200, but only under a tightly matched serving test: one 8,000-token prompt, 1,000-token response, and 50-token-per-second-per-user target. FP4 and different serving stacks are part of the result, so it is not a general agent-performance verdict. GB200 is cheaper on listed token cost ($0.041 versus...

  1. 01

    GB300 led B300, B200 and MI355X, which recorded 12,253, 11,606 and 7,692 tokens per second per GPU.

  2. 02

    InferenceX excluded platforms without measurements at the 50-token-per-second-per-user operating point.

  3. 03

    GB300 and GB200 used Dynamo vLLM; B300, B200 and MI355X used vLLM, all at FP4.

MiniMax pitches M2.7 for complex agent harnesses and software-engineering work. InferenceX’s new ranking answers a more contained question: on a single-turn chat workload, GB300 NVL72 delivered the highest measured throughput at 13,707 tokens per second per GPU.

GB200 NVL72 came second at 12,515 tokens per second per GPU, a 10% gap. B300, B200 and MI355X followed at 12,253, 11,606 and 7,692 tokens per second per GPU, respectively.

A shared responsiveness constraint

Each platform was read at a matched target of 50 tokens per second per user. The test used one chat turn with 8,000 input tokens and 1,000 output tokens. InferenceX excludes hardware without a measurement at that operating point, rather than comparing a system at a different user-speed tradeoff.

The result includes the serving stack

This is not a hardware-only comparison. All listed measurements use FP4 precision, while GB300 and GB200 use Dynamo vLLM; B300, B200 and MI355X use vLLM. InferenceX says it measures real hardware, sweeps concurrency to trace each platform’s throughput-versus-interactivity frontier, and reruns tests as engines and configurations change.

A speed result, not an agent verdict

MiniMax announced M2.7 in March and describes it as supporting complex agent harnesses, agent teams, dynamic tool search and software-engineering tasks. Those are broader patterns than the ranking’s one prompt-and-response exchange. The benchmark therefore establishes a leader for its specified serving case, not across every workload MiniMax associates with M2.7.

Speed and listed token cost also point in different directions. InferenceX puts GB300 at $0.047 per million tokens, versus $0.041 for GB200, using measured throughput and GPU-hour rates from SemiAnalysis’s AI Cloud total-cost-of-ownership model. The page reruns continuously, but its newest input result was dated June 13, so later engine releases or configurations could change the order.

Sources

  1. inferencex.semianalysis.comFastest GPU for MiniMax M2.7 Inference: Live Rankings | InferenceX
  2. minimax.ioMiniMax M2.7: Early Echoes of Self-Evolution