Benchmark Methodology

How a VoxelBench benchmark runs, how to prepare a server so that its results are reliable and comparable, and which factors move the score. The score itself is explained in How Scoring Works.

The Benchmark Process

A full benchmark is started with /bench start by a player connected to the server. The plugin runs 16 tests in three phases, and 14 of them make the score (see How Scoring Works):

  1. Hardware โ€” disk, network and memory, while the JVM warms up
  2. Gameplay โ€” Minecraft workloads: chunkLoading, hopper, worldSave, explosion, redstone, blockPhysics, chunkTicking, lightingUpdate, tickingTileEntity, mobAI, boneMealGrowth
  3. CPU โ€” singleCoreBenchmark then multiCore, last, once the JVM is fully warmed up

The plugin does the rest on its own: it spawns the entities, loads the chunks, drives the redstone, sets off the explosions, then cleans up and sends the report. In detail:

  • Before the run, it checks that the command comes from a connected player, that no other test is running and that the local cooldown between benchmarks (30 minutes by default) has passed. Pre-flight checks then look for risks โ€” a target world that is not flat or not a voxelbench_* world, other players online, a server already under load (TPS under 19 or MSPT over 30 ms), free hosting, plugins likely to interfere โ€” and list them on a screen where you confirm or cancel.
  • The JVM warms up for 30 seconds before the first test.
  • Every test runs with fixed parameters in the standard mode (benchmark-mode: standard, the default), so every server runs the same workload.
  • Gameplay tests ignore their first 5 seconds of samples, and the plugin pauses between two tests until the previous one's chunks are unloaded.
  • A test that is skipped or fails does not stop the run, but the report then gets no score.
  • The run takes place in a benchmark world: the pinned world if you set one, otherwise a temporary flat world (voxelbench_temp_โ€ฆ) created for the run and deleted afterwards. On Folia, which cannot create a world while running, pin one.

The steps as seen from the server are in Benchmarks, and the worlds in Benchmark Worlds.

Test Duration

A full benchmark usually takes 10 to 15 minutes. Faster servers complete the tests quicker โ€” this is itself a performance indicator, and the total run time, warm-up included, is factored into the score: a run of 5 minutes or less earns ร—1.15, the multiplier is neutral around 13 minutes, and falls to ร—0.90 at 20 minutes and ร—0.70 at 40 (see Duration). Each test also has a time limit: a test that stops making progress or runs past its ceiling is ended, and counts as failed.

Several Runs

/bench start <runs> chains 1 to 20 benchmarks, 10 seconds apart, each with its own report; /bench start warmup adds a first run whose results are discarded. Several runs show how much your results vary from one run to the next (see Multi-Run Mode).

Preparing for a Reliable Benchmark

Getting consistent, comparable results requires controlling the test environment. Here's a checklist.

Server Configuration

Minecraft Version

  • Use a stable release (not snapshots or pre-releases)
  • Compare results made on the same Minecraft version
  • Newer versions may have different performance characteristics

Server Software

  • Paper is recommended for the most optimized results
  • Spigot, Purpur and Folia are supported, and hybrid servers are detected with some limits (see Compatibility)
  • Results from different server software do not compare directly

Java Version

  • The Minecraft version sets the minimum: Java 16 or 17 for 1.17, 17 for 1.18 to 1.20.4, 21 for 1.20.5 to 1.21, 25 for 26.x (see Java Requirements)
  • Use the newest Java your Minecraft version supports, and the same one when comparing

JVM Arguments (Critical)

The garbage collector configuration has a major impact on results: its pauses stop the main thread, and show up as slow ticks in the 95th and 99th percentiles of tick time, which weigh the most in the score.

  • Aikar's flags are the recommended baseline for most servers
  • G1GC (default with Aikar's flags) is a good balance of throughput and latency
  • ZGC (-XX:+UseZGC) produces very short pauses but may use more memory
  • Avoid default JVM flags without tuning โ€” they lead to long GC pauses

Key JVM settings to verify:

  • -Xms and -Xmx should be equal (prevents heap resizing during tests)
  • Give the server enough memory for the heaviest tests: before each test, the plugin estimates its memory peak and skips a test that could exhaust the heap, and a report with a skipped test gets no score
  • Don't over-allocate: a heap far larger than needed can lengthen GC pauses

Environment

No Other Players

  • Run benchmarks with no player online besides the one who started it; the pre-flight checks warn about other players
  • Player activity creates unpredictable load that varies between runs
  • Even AFK players generate chunk ticks and entity interactions
  • Stay connected until the end: some tests need you, and are skipped if you leave

No Other Plugins (Ideal)

  • For the most accurate hardware comparison, test with only the VoxelBench plugin
  • If you need to test your production configuration, understand that other plugins add overhead
  • The pre-flight checks name the plugins most likely to interfere (world protection such as WorldGuard, GriefPrevention, Towny or Lands, profilers such as spark, and Essentials)

Dedicated Resources

  • Close other applications on the machine (or VM/container)
  • On shared hosting: other customers' servers affect your results
  • CPU throttling (power saving modes, thermal throttling) will produce inconsistent results
  • If possible, run the benchmark during low-activity hours on the host machine

World State

  • Let the benchmark run in a benchmark world: a pinned voxelbench_* world (/bench world create, then /bench world set) or the temporary flat world
  • Your own map then plays no part in the results: the terrain is flat and the test zones are built for the run

Network

  • The latency between your server and voxelbench.com does not affect the score: all measurements are taken on the server
  • The network test measures how fast the server serializes game data (NBT), not your connection, and carries no weight
  • The connection must only be stable enough to send the report

What Makes Results Comparable

Two benchmarks are directly comparable when:

FactorMust MatchWhy
Scoring versionYesScores from two versions of the rules, or from a standard and a stress limit run, are not comparable
Minecraft versionYesPerformance varies significantly between versions
Server softwareYesPaper and Spigot have different optimization levels
Java versionIdeallyA newer Java can change tick times and GC behaviour
JVM flags / GCIdeallyGC strategy affects pause times and MSPT stability
RAM allocationSimilarToo little and tests are skipped; too much and GC pauses can grow
Players onlineYes (only the one running it)Players add variable load
Plugin setIdeallyEach plugin adds baseline overhead

A single run is only one sample: compare several runs of each server before concluding.

Hosting Provider Comparison

The Hosting pages and the Offers catalogue show, for each offer, its strongest evidence. The strongest is a certified result: VoxelBench buys the offer like any customer, benchmarks it three times โ€” auto-bench runs at least an hour apart, or runs a moderator selects โ€” and keeps the median run, checks that the machine matches the advertised specifications (a mismatch ends in a rejection), and publishes the report once the provider has approved it. The provider may contest a result before approving it. See Certification and Comparing Hosting Providers.

This way the workload is the same as on your own server โ€” the same standard tests with the same parameters, in a flat benchmark world โ€” and what changes is the hosting infrastructure itself: CPU, memory, disk I/O, and how well resources are isolated from other customers.

Factors That Hurt Your Score

FactorImpactHow to Fix
Slow ticks (MSPT over 50 ms)Tick scores fall steeply past the 50 ms budgetFaster single-thread CPU, fewer plugins, tuned configs
Lag spikes and long GC pausesThe slowest ticks (95th and 99th percentiles) weigh the mostTune the JVM: G1GC with Aikar's flags, or ZGC
Unstable TPS or MSPTStability multiplier down to about ร—0.87Stop background tasks, avoid CPU throttling
A skipped or failed testNo score at allEnough memory, a connected player until the end, a benchmark world
Unusable test dataโˆ’8 % per test, up to โˆ’40 %Let the tests run on an idle server
Slow run (over about 13 minutes)Duration multiplier below ร—1.00: ร—0.90 at 20 minutes, ร—0.70 at 40Better hardware
One weak categorySynergy follows the weakest categoryDon't neglect memory or disk
Slow disk I/OReduces hardware scoreSSD rather than HDD, NVMe rather than SATA
Low memory bandwidth or high latencyReduces hardware scoreFaster RAM, proper DIMM configuration

Factors That Don't Affect Your Score

FactorWhy Not
Network latency / pingAll measurements are taken on the server
Your map and buildsThe tests run in a flat benchmark world
Time since the server startedThe plugin warms the JVM up for 30 seconds before the first test
Time of day (unless the machine is shared)Only other customers' load changes with the hour