Benchmark Methodology
How a VoxelBench benchmark runs, how to prepare a server so that its results are reliable and comparable, and which factors move the score. The score itself is explained in How Scoring Works.
The Benchmark Process
A full benchmark is started with /bench start by a player connected to the server. The plugin runs 16 tests in three phases, and 14 of them make the score (see How Scoring Works):
- Hardware โ
disk,networkandmemory, while the JVM warms up - Gameplay โ Minecraft workloads:
chunkLoading,hopper,worldSave,explosion,redstone,blockPhysics,chunkTicking,lightingUpdate,tickingTileEntity,mobAI,boneMealGrowth - CPU โ
singleCoreBenchmarkthenmultiCore, last, once the JVM is fully warmed up
The plugin does the rest on its own: it spawns the entities, loads the chunks, drives the redstone, sets off the explosions, then cleans up and sends the report. In detail:
- Before the run, it checks that the command comes from a connected player, that no other test is running and that the local cooldown between benchmarks (30 minutes by default) has passed. Pre-flight checks then look for risks โ a target world that is not flat or not a
voxelbench_*world, other players online, a server already under load (TPS under 19 or MSPT over 30 ms), free hosting, plugins likely to interfere โ and list them on a screen where you confirm or cancel. - The JVM warms up for 30 seconds before the first test.
- Every test runs with fixed parameters in the standard mode (
benchmark-mode: standard, the default), so every server runs the same workload. - Gameplay tests ignore their first 5 seconds of samples, and the plugin pauses between two tests until the previous one's chunks are unloaded.
- A test that is skipped or fails does not stop the run, but the report then gets no score.
- The run takes place in a benchmark world: the pinned world if you set one, otherwise a temporary flat world (
voxelbench_temp_โฆ) created for the run and deleted afterwards. On Folia, which cannot create a world while running, pin one.
The steps as seen from the server are in Benchmarks, and the worlds in Benchmark Worlds.
Test Duration
A full benchmark usually takes 10 to 15 minutes. Faster servers complete the tests quicker โ this is itself a performance indicator, and the total run time, warm-up included, is factored into the score: a run of 5 minutes or less earns ร1.15, the multiplier is neutral around 13 minutes, and falls to ร0.90 at 20 minutes and ร0.70 at 40 (see Duration). Each test also has a time limit: a test that stops making progress or runs past its ceiling is ended, and counts as failed.
Several Runs
/bench start <runs> chains 1 to 20 benchmarks, 10 seconds apart, each with its own report; /bench start warmup adds a first run whose results are discarded. Several runs show how much your results vary from one run to the next (see Multi-Run Mode).
Preparing for a Reliable Benchmark
Getting consistent, comparable results requires controlling the test environment. Here's a checklist.
Server Configuration
Minecraft Version
- Use a stable release (not snapshots or pre-releases)
- Compare results made on the same Minecraft version
- Newer versions may have different performance characteristics
Server Software
- Paper is recommended for the most optimized results
- Spigot, Purpur and Folia are supported, and hybrid servers are detected with some limits (see Compatibility)
- Results from different server software do not compare directly
Java Version
- The Minecraft version sets the minimum: Java 16 or 17 for 1.17, 17 for 1.18 to 1.20.4, 21 for 1.20.5 to 1.21, 25 for 26.x (see Java Requirements)
- Use the newest Java your Minecraft version supports, and the same one when comparing
JVM Arguments (Critical)
The garbage collector configuration has a major impact on results: its pauses stop the main thread, and show up as slow ticks in the 95th and 99th percentiles of tick time, which weigh the most in the score.
- Aikar's flags are the recommended baseline for most servers
- G1GC (default with Aikar's flags) is a good balance of throughput and latency
- ZGC (
-XX:+UseZGC) produces very short pauses but may use more memory - Avoid default JVM flags without tuning โ they lead to long GC pauses
Key JVM settings to verify:
-Xmsand-Xmxshould be equal (prevents heap resizing during tests)- Give the server enough memory for the heaviest tests: before each test, the plugin estimates its memory peak and skips a test that could exhaust the heap, and a report with a skipped test gets no score
- Don't over-allocate: a heap far larger than needed can lengthen GC pauses
Environment
No Other Players
- Run benchmarks with no player online besides the one who started it; the pre-flight checks warn about other players
- Player activity creates unpredictable load that varies between runs
- Even AFK players generate chunk ticks and entity interactions
- Stay connected until the end: some tests need you, and are skipped if you leave
No Other Plugins (Ideal)
- For the most accurate hardware comparison, test with only the VoxelBench plugin
- If you need to test your production configuration, understand that other plugins add overhead
- The pre-flight checks name the plugins most likely to interfere (world protection such as WorldGuard, GriefPrevention, Towny or Lands, profilers such as spark, and Essentials)
Dedicated Resources
- Close other applications on the machine (or VM/container)
- On shared hosting: other customers' servers affect your results
- CPU throttling (power saving modes, thermal throttling) will produce inconsistent results
- If possible, run the benchmark during low-activity hours on the host machine
World State
- Let the benchmark run in a benchmark world: a pinned
voxelbench_*world (/bench world create, then/bench world set) or the temporary flat world - Your own map then plays no part in the results: the terrain is flat and the test zones are built for the run
Network
- The latency between your server and voxelbench.com does not affect the score: all measurements are taken on the server
- The
networktest measures how fast the server serializes game data (NBT), not your connection, and carries no weight - The connection must only be stable enough to send the report
What Makes Results Comparable
Two benchmarks are directly comparable when:
| Factor | Must Match | Why |
|---|---|---|
| Scoring version | Yes | Scores from two versions of the rules, or from a standard and a stress limit run, are not comparable |
| Minecraft version | Yes | Performance varies significantly between versions |
| Server software | Yes | Paper and Spigot have different optimization levels |
| Java version | Ideally | A newer Java can change tick times and GC behaviour |
| JVM flags / GC | Ideally | GC strategy affects pause times and MSPT stability |
| RAM allocation | Similar | Too little and tests are skipped; too much and GC pauses can grow |
| Players online | Yes (only the one running it) | Players add variable load |
| Plugin set | Ideally | Each plugin adds baseline overhead |
A single run is only one sample: compare several runs of each server before concluding.
Hosting Provider Comparison
The Hosting pages and the Offers catalogue show, for each offer, its strongest evidence. The strongest is a certified result: VoxelBench buys the offer like any customer, benchmarks it three times โ auto-bench runs at least an hour apart, or runs a moderator selects โ and keeps the median run, checks that the machine matches the advertised specifications (a mismatch ends in a rejection), and publishes the report once the provider has approved it. The provider may contest a result before approving it. See Certification and Comparing Hosting Providers.
This way the workload is the same as on your own server โ the same standard tests with the same parameters, in a flat benchmark world โ and what changes is the hosting infrastructure itself: CPU, memory, disk I/O, and how well resources are isolated from other customers.
Factors That Hurt Your Score
| Factor | Impact | How to Fix |
|---|---|---|
| Slow ticks (MSPT over 50 ms) | Tick scores fall steeply past the 50 ms budget | Faster single-thread CPU, fewer plugins, tuned configs |
| Lag spikes and long GC pauses | The slowest ticks (95th and 99th percentiles) weigh the most | Tune the JVM: G1GC with Aikar's flags, or ZGC |
| Unstable TPS or MSPT | Stability multiplier down to about ร0.87 | Stop background tasks, avoid CPU throttling |
| A skipped or failed test | No score at all | Enough memory, a connected player until the end, a benchmark world |
| Unusable test data | โ8 % per test, up to โ40 % | Let the tests run on an idle server |
| Slow run (over about 13 minutes) | Duration multiplier below ร1.00: ร0.90 at 20 minutes, ร0.70 at 40 | Better hardware |
| One weak category | Synergy follows the weakest category | Don't neglect memory or disk |
| Slow disk I/O | Reduces hardware score | SSD rather than HDD, NVMe rather than SATA |
| Low memory bandwidth or high latency | Reduces hardware score | Faster RAM, proper DIMM configuration |
Factors That Don't Affect Your Score
| Factor | Why Not |
|---|---|
| Network latency / ping | All measurements are taken on the server |
| Your map and builds | The tests run in a flat benchmark world |
| Time since the server started | The plugin warms the JVM up for 30 seconds before the first test |
| Time of day (unless the machine is shared) | Only other customers' load changes with the hour |