How Scoring Works

voxelbench.com turns the results of a full benchmark into one score and a rank. This page explains which tests count, how each category is scored, what adjusts the total and how the rank is decided, following the current rules, Standard scoring 8.0.0.

Where the Score Comes From

The plugin runs the tests and sends their raw results; voxelbench.com computes the score, and the plugin only shows what the site returns (see Benchmarks). The rules have a version, shown on every report as Standard scoring 8.0.0. When the rules change, a report scored with an earlier version is marked earlier scoring: scores from two versions are not directly comparable.

Stress limit runs use a separate scale, Stress-limit scoring, with its own version and a total that tops out around 25,000: never compare it with a standard score. Custom profile runs are kept without a score.

Which Tests Count

A full benchmark (/bench start) runs 16 tests. 14 of them make the score, in three categories:

CategoryWeightTests
Single-Core40 %redstone, blockPhysics
Gameplay40 %Exploration: chunkLoading, chunkTicking, lightingUpdate. Entities: mobAI. Mechanics: hopper, explosion, tickingTileEntity, worldSave, boneMealGrowth
Hardware20 %memory, disk, multiCore
  • network must be in the report, but carries no weight.
  • singleCoreBenchmark, run at the end of the benchmark, is kept in the report but not scored.
  • The other tests of the catalogue (mobSpawn, villagerTrading, liquidPhysics…) are only run by /bench test or custom profiles, and never count.

The plugin's catalogue lists redstone and blockPhysics among the gameplay tests: for the score, they make the single-core category, because both put all their load on the server's main thread.

Single-core and gameplay together make up 80 % of the score because Minecraft runs its game loop on one thread. The fastest disk in the world won't help if the CPU can't keep up with tick processing.

A Missing, Failed or Skipped Test

A report is scored only when every required test (the 14 above, plus network) is present, passed, and was not skipped. A skipped test β€” for example one the plugin refused to start because the server's memory was too short β€” is not a fault of the server, but it is not a measurement either: the report is kept, without score or rank, and stays off the leaderboard. See Unit Tests for the reasons a test can be skipped or fail.

A test that ran but produced unusable data (no samples, no work done, entities that never ticked, zones that could not tick…) costs 8 % of the total, cumulated per test, up to βˆ’40 %.

Base Score

Each category gets its own score, then the three are combined with their weights (40 / 40 / 20). The combination is mostly a weighted average (85 %), with a geometric-mean part (15 %): a very weak category pulls the total down more than a plain average would.

Single-Core Performance (40 %)

redstone (53 %) drives pistons with redstone circuits and blockPhysics (47 %) makes blocks fall: both measure how well the main thread keeps its 20 ticks per second.

What Matters Most

  • MSPT (milliseconds per tick) counts for 75 % of a test's tick score, TPS for 25 %. Unlike TPS, which caps at 20, MSPT shows the margin left: 20 TPS at 10 ms leaves 40 ms of headroom, 20 TPS at 48 ms is on the edge.
  • Worst ticks weigh most. The MSPT part is 40 % the 95th percentile, 40 % the 99th and only 20 % the average; the TPS part mixes the median (35 %), the 5th percentile (30 %), the 1st percentile (15 %) and the average (20 %).
  • The curve. Below 20 ms per tick, a test gets the best value; between 20 and 50 ms it falls along an S-curve centred on 35 ms; beyond the 50 ms budget it falls steeply, with an extra penalty past 100 ms. On the TPS side, an average of 19.8 or more gets the full value, and a minimum TPS under 18 costs extra.
  • Stability is rewarded: low variation of TPS and MSPT, and worst ticks close to the average, each add a bonus, and so does a low average MSPT.

The two tests also earn a bonus together: when their scores agree (up to +8 %), and when both hold at least 19.0, 19.5 or 19.8 TPS on average (+5, +10 or +15 %). Very different scores between them cost 5 %.

Gameplay Score (40 %)

Real Minecraft workloads, built by the plugin in the benchmark world.

Sub-categories

Exploration (30 % of gameplay)

  • chunkLoading (43 %): chunks generated and loaded per second, the load time, and the tick performance meanwhile
  • chunkTicking (33 %): random ticks at a raised tick speed, scored on tick performance
  • lightingUpdate (24 %): light engine updates, scored on tick performance

Entities (25 % of gameplay)

  • mobAI alone: a village of 2,400 trading villagers, then an invasion of 1,200 hostile mobs, scored on tick performance and on how close the two phases stay

Mechanics (45 % of gameplay) β€” the heaviest gameplay sub-category

  • hopper (21 %): items moved per second (85 %) and tick performance (15 %)
  • explosion (23 %): blocks destroyed per second, with a penalty when the minimum TPS falls under 15
  • tickingTileEntity (19 %): furnaces, hoppers and spawners ticking, scored on tick performance
  • worldSave (17 %): the impact of a full world save on the tick. Its write throughput no longer counts, so this test weighs little and scores nearly the same everywhere
  • boneMealGrowth (20 %): the growth rate of saplings and crops (75 %) and tick performance (25 %)

A mechanics test whose measurement the plugin flagged as less reliable cannot score above the average of the other mechanics tests.

Why Mechanics Are Weighted Highest

Mechanical interactions are what players actually build on survival and technical servers. A server that handles chunk loading well but chokes on hopper chains isn't useful for a typical SMP or technical server.

Balance

When the three sub-scores are within 25 % of each other, gameplay gets +3 %; when the weakest is under 40 % of the strongest, βˆ’2 %.

Hardware Score (20 %)

Raw system capabilities, independent of Minecraft load:

Memory (50 % of hardware) β€” the most impactful hardware measurement for Minecraft

  • Sequential read and write, random access, copy throughput, and memory latency
  • Random access weighs the most, because Minecraft's world data is mostly read in random order; latency also scales the whole memory score

Disk (25 %)

  • Sequential throughput, and random 4K throughput measured with direct I/O (bypassing the system cache); random access weighs more
  • A penalty when the 99th-percentile latency goes over 1 ms, and a bonus of up to +10 % for a high number of operations per second
  • Matters for world saves and chunk loading from disk

Multi-Core (25 %)

  • Parallel throughput across the cores the server may use, and how well it scales
  • A bonus for 4, 8 and 16 cores or more (+3, +7 and +13 %): useful for Paper's asynchronous work, garbage collection and plugins

Why Hardware Is Only 20 %

Because Minecraft ticks on one thread. A server with a blazing fast NVMe and 128 GB of RAM but a weak single-core CPU will still have bad TPS. Hardware matters, but single-thread performance dominates.

Multipliers

After the base score, several multipliers adjust it. The report page shows four of them: Synergy, Stability, Duration and GC.

Synergy

Depends on the weakest of the three categories only: the stronger your weakest category, the higher the bonus, from Γ—0.90 to Γ—1.15. Improving any category can therefore never lower the score.

Duration

The total run time, JVM warm-up included. A faster server finishes sooner, which is itself a performance indicator:

Run timeMultiplier
5 minutes or lessΓ—1.15
5 to 10 minutesΓ—1.15 down to Γ—1.05
10 to 20 minutesΓ—1.05 down to Γ—0.90 (neutral at about 13 min 20 s)
20 to 40 minutesΓ—0.90 down to Γ—0.70
Beyond 40 minutesβˆ’0.03 per extra minute, down to Γ—0.50

Stability

Consistency across the benchmark: the average variation of TPS and of MSPT over six tests (redstone, blockPhysics, chunkLoading, chunkTicking, mobAI, hopper), from about Γ—0.87 to Γ—1.065. Inconsistent results suggest external interference: other processes, CPU throttling, noisy neighbours on shared hosting.

Garbage Collection

The rules include a garbage-collection multiplier (overhead, long pauses, pause frequency), but the plugin does not send the measurement it reads today, so it stays at Γ—1.00. GC pauses still count: they show up as slow ticks in the MSPT percentiles of every test.

Invalid Data

The βˆ’8 % per unusable test described above, down to Γ—0.60.

The product then goes through a gentle curve centred on 250,000, the median expected score: it slightly compresses scores above it and stretches those below. A fixed Γ—0.98 factor, left from an older cross-validation step, is also applied. There is no maximum score.

Rank Badges

The rank compares your score with the other reports: public and certified standard reports of the last 6 months, certified ones counting four times.

RankLabelPosition
SSLegendaryTop 1 %
SExceptionalTop 5 %
AExcellentTop 15 %
BVery GoodTop 35 %
CGoodTop 60 %
DAverageTop 85 %
EBelow AverageTop 95 %
FPoorThe rest

The thresholds are recomputed every 6 hours, and every rank with them: a report's rank can change over time while its score stays the same. Until the pool holds 10 reports, fixed thresholds apply instead (SS from 450,000, then 400,000, 350,000, 300,000, 250,000, 200,000 and 150,000 for E).

Stress limit runs have fixed thresholds: SS from 25,000, S from 20,000, A from 16,000, B from 12,000, C from 9,000, D from 6,000, E from 3,000.

Key Takeaways

  1. MSPT matters more than TPS β€” TPS caps at 20, MSPT shows the real margin, and it weighs three times more
  2. Worst-case matters more than average β€” the 95th and 99th percentiles make most of the tick score
  3. A complete run is required β€” one skipped or failed required test, and the report gets no score
  4. Stability is rewarded β€” Consistent 18 TPS beats erratic 20-then-12 TPS
  5. Real gameplay dominates β€” 80 % of the score depends on the main thread and Minecraft workloads
  6. The weakest category sets the synergy β€” a well-rounded server outscores one with a single strength
  7. The rank is relative β€” it places your score among recent public reports