Download VoxelBench
Get the VoxelBench plugin for your Minecraft server
a0759bc184c1858427dfbc7c6874eeaf61ad0db0a3e4a98906ac5ea0e6714936Benchmarks no longer remove anything outside their own worlds
VoxelBench 2.0.3 fixes cleanups that could delete what players had built or tamed in your worlds. It also corrects how three tests are measured on Folia and how several tests count their load in custom profiles, and brings a batch of smaller fixes. If you run 2.0.2 or earlier, read the first section before your next benchmark.
If you run 2.0.2 or earlier
Up to 2.0.2, several cleanups removed things VoxelBench had not placed, in worlds that were not its own.
- A multi-run could delete regions of your main world.
/bench startwith more than one run, or withwarmup, and no pinned world — the default setup — prepared its test zones in the main world, while the tests themselves ran in a temporary benchmark world. Between two runs it then reset those regions of the main world: chunks unloaded without saving, and the region, entity and POI files around each zone deleted. The zones are drawn 1,000 to 10,000 blocks from spawn, and anything built there was lost. Pinning the main world with/bench world sethad the same effect. Auto-bench jobs with several runs used the same reset, and reached the main world whenever they fell back on it: when no temporary world could be created (Folia refuses to create one while the server runs), or when the main world was pinned. - Every run removed animals and other entities from every loaded world, on
Paper and Spigot. At the start of each
/bench start, custom profile, tier and stress-limit run, auto-bench jobs included, VoxelBench removed from every loaded world all loaded villagers and iron golems, farm animals, wolves, cats, horses and other mounts — tamed or named ones included —, hostile mobs, dropped items, projectiles, minecarts and boats (chest and hopper variants included). Folia skipped this sweep. /bench zones cleanand/bench stopcould sweep the main world. Without a pinned world,/bench zonesplaced its zone at x = 50,000, z = 50,000 of the main world, andcleanremoved every entity except players, paintings and item frames in a 400-block-wide box around it./bench stop, a shutdown during a run and the recovery after a crash removed test-like entities (villagers, cats, golems, zombies, skeletons, creepers, items, TNT…) around fixed test coordinates between x = 50,000 and 70,000, in the test's world — or in the main world when that world was unknown.
Update before your next multi-run. Until you can, pin a voxelbench_*
world first — /bench world create bench, then
/bench world set voxelbench_bench — so that pre-warm, region resets and the
sweeps of /bench zones clean and /bench stop target that world. Pinning
does not stop the start-of-run sweep on Paper and Spigot: only the update
does. On Folia, where a world cannot be created while the server runs, avoid
multi-runs until you update.
What 2.0.3 changes
Outside voxelbench_* worlds — its temporary benchmark worlds and the ones
made with /bench world create — VoxelBench no longer removes anything it did
not place itself.
- A multi-run takes its world before pre-warming, so pre-warm, shared zones
and region resets all target the world the tests run in. Region files are
reset, and chunks force-loaded, only in a
voxelbench_*world; anywhere else that step is skipped with a message and the run goes on. Auto-bench multi-runs follow the same rule. - The start-of-run sweep only removes entities in
voxelbench_*worlds. Weather and time are still set in every world during a run, then restored. /bench zonesshows the zones of the last run, in that run's world, and says so when that world is gone;cleanrefuses zones outside avoxelbench_*world./bench stop, a shutdown and crash recovery sweep onlyvoxelbench_*worlds, with no fallback on the main world.
Tests still remove what they built themselves, wherever they were allowed to run. Two consequences you may notice:
- A server whose worlds hold many loaded entities keeps them during the run, as in normal play.
- A multi-run in a world that is not a
voxelbench_*world — the main world pinned, or Folia without a pinnedvoxelbench_*world — no longer resets regions between runs: it says so, and later runs reuse the chunks already generated there.
Measurements
Folia: three tests measured where they work
On Folia, a test's headline TPS and MSPT come from the regions it works in.
lightingUpdate and tickingTileEntity build at their own coordinates, but
were measured on the run's shared zones — where they build nothing — as soon
as they used more than one zone, the standard benchmark included.
playerWorldLoad was measured at the world centre only, while its simulated
players spread up to 8,000 blocks away. Those figures described regions at
rest, at about 20 TPS, whatever the load.
They are now measured in the regions each test builds in, and
lightingUpdate's Folia measurement starts once its chunks are loaded, as on
Paper. Folia results for these three tests may come out lower than in
2.0.2: they now reflect the load the test puts on the server.
lightingUpdate and tickingTileEntity count in the site's Gameplay score,
so a standard Folia score can move with them. Paper and Spigot results are
unchanged.
Custom profiles: loads counted on the zones actually used
In a custom profile, a step whose dispersedZones is above 1 always runs on
the run's 8 zones. Four tests computed their totals on the step's value
instead, so a step at 2 to 7 was measured on the wrong basis:
- hopper (Paper and Spigot) filled only the zones the value named — 2 of
8 in
free-host, 4 of 8 inlow-memoryandexample— and the other structures ran empty; a value of 9 or 10 made the test fail. - explosion counted
tntCount× the value as requested, while placing charges in all 8 zones:free-hostplaced 160 charges for 40 "requested". The detonation ratio could reach 4, and the share of charges counted as presented could fall to 0, which lowered the site's confidence in the result. - mobSpawn could publish a survival rate above 100 %. No bundled profile was affected.
- chunkLoading divided its memory cap (30 % of the heap) by the value rather than by the zones it walks: where that cap applied, a step at 2 could load up to four times the chunks the cap allowed.
These totals are now counted on the zones each test actually uses. The
standard benchmark, stress-limit runs, single /bench test runs and steps at
1 or 8 are unchanged. Results of the affected profiles show a step in their
history: compare 2.0.3 runs with 2.0.3 runs.
Stress profiles: plateauBoost now applies
stress.ramp.plateauBoost was read and published in reports, but never
applied: the anti-plateau boost always ran. It now does what the profile says.
Only profiles that set it to false change — among the bundled ones,
stress-freehost, which now ramps without the boost, as intended for small
hosts. /bench stresslimit and the other bundled stress profiles keep their
ramp.
Bundled profiles say what they do
Six bundled profiles (example, free-host, low-memory, showcase,
stress-freehost, stress-monster) now describe what they really do: 8 zones
for any step above 1, except lightingUpdate and tickingTileEntity, which
build exactly as many zones as the step says; the config.yml limits that
silently cap chunk loading, mob spawn and explosion steps; six stress types
of seven in stress-monster; free-host in English. A copy you never edited
is updated automatically; an edited copy is left as it is.
A profile's description is part of its fingerprint, the profile hash that
/bench custom info shows and that voxelbench.com uses to group runs of the
same profile. free-host, stress-freehost and stress-monster have a new
description, so their results from 2.0.3 on are no longer grouped with
earlier ones. The other three changed only in comments and keep their
fingerprint.
Permissions
- A
kind: stresslimitprofile now requiresvoxelbench.stresslimit, like/bench stresslimit;voxelbench.customalone was enough to run one. Stress profiles are offered — in tab completion, on the profile screen and in the list of available profiles — only to those who can run them. - The profile screen requires
voxelbench.customorvoxelbench.start(voxelbench.guialone opened it), and its refresh button runs/bench custom reload, which requiresvoxelbench.reload.
Smaller fixes
- World messages that work when you follow them. A leftover world folder
no longer points you to
/bench world delete, which cannot remove a world that is not loaded; the pinning advice namesvoxelbench_<name>, the name/bench world creategives; the Multiverse message says the world was handed to/mv importand asks you to check/mv list. /bench world listno longer has a[MV]marker: it never appeared./bench verifytells apart a code made for another server, an expired code and too many attempts, instead of "server error, try again later"; a server that is already verified is told so.- Remote monitoring over the plan's server limit now pauses, checks again every 10 minutes and says to turn monitoring off for another server or change plan. It used to retry every minute with a bare HTTP 403.
- A second
/bench linkwhile one is waiting for confirmation is refused. It used to open a second link in parallel, one token overwriting the other. /api/metrics:ram_usedandram_freeare percentages of the maximum heap, and theirunitnow says%instead ofMB. Values are unchanged./bench monitor web startandstatusshow the real address —httpsandmonitor.https.portwhen HTTPS is on.- Chunks VoxelBench force-loads are now released on
/bench stop, on shutdown and after a crash, in every world. A chunk forced when the server crashed used to stay force-loaded for good. - Score screens, chat and the DiscordSRV message use the site's weights — Single-Core 40 %, Hardware 20 %, Gameplay 40 % (they said 45/20/35). They name Redstone and Block Physics under Single-Core, as the site does, and mark the network and single-core CPU tests as not counted in the site's score. VoxelBench shows the site's score; it computes none.
- The
config.ymlcomment on how long voxelbench.com keeps reports now matches the site (new installs only).
Upgrade
Drop the new jar in plugins/ and restart. No config or data migration.
Bundled profiles you never edited are refreshed; edited ones are left alone.
A group that has voxelbench.custom without voxelbench.stresslimit can no
longer run stress profiles: grant the node if it should. A tool reading the
unit of ram_used or ram_free in /api/metrics now gets %.
Benchmarks no longer remove anything outside their own worlds
VoxelBench 2.0.3 fixes cleanups that could delete what players had built or tamed in your worlds. It also corrects how three tests are measured on Folia and how several tests count their load in custom profiles, and brings a batch of smaller fixes. If you run 2.0.2 or earlier, read the first section before your next benchmark.
If you run 2.0.2 or earlier
Up to 2.0.2, several cleanups removed things VoxelBench had not placed, in worlds that were not its own.
- A multi-run could delete regions of your main world.
/bench startwith more than one run, or withwarmup, and no pinned world — the default setup — prepared its test zones in the main world, while the tests themselves ran in a temporary benchmark world. Between two runs it then reset those regions of the main world: chunks unloaded without saving, and the region, entity and POI files around each zone deleted. The zones are drawn 1,000 to 10,000 blocks from spawn, and anything built there was lost. Pinning the main world with/bench world sethad the same effect. Auto-bench jobs with several runs used the same reset, and reached the main world whenever they fell back on it: when no temporary world could be created (Folia refuses to create one while the server runs), or when the main world was pinned. - Every run removed animals and other entities from every loaded world, on
Paper and Spigot. At the start of each
/bench start, custom profile, tier and stress-limit run, auto-bench jobs included, VoxelBench removed from every loaded world all loaded villagers and iron golems, farm animals, wolves, cats, horses and other mounts — tamed or named ones included —, hostile mobs, dropped items, projectiles, minecarts and boats (chest and hopper variants included). Folia skipped this sweep. /bench zones cleanand/bench stopcould sweep the main world. Without a pinned world,/bench zonesplaced its zone at x = 50,000, z = 50,000 of the main world, andcleanremoved every entity except players, paintings and item frames in a 400-block-wide box around it./bench stop, a shutdown during a run and the recovery after a crash removed test-like entities (villagers, cats, golems, zombies, skeletons, creepers, items, TNT…) around fixed test coordinates between x = 50,000 and 70,000, in the test's world — or in the main world when that world was unknown.
Update before your next multi-run. Until you can, pin a
voxelbench_*world first —/bench world create bench, then/bench world set voxelbench_bench— so that pre-warm, region resets and the sweeps of/bench zones cleanand/bench stoptarget that world. Pinning does not stop the start-of-run sweep on Paper and Spigot: only the update does. On Folia, where a world cannot be created while the server runs, avoid multi-runs until you update.What 2.0.3 changes
Outside
voxelbench_*worlds — its temporary benchmark worlds and the ones made with/bench world create— VoxelBench no longer removes anything it did not place itself.- A multi-run takes its world before pre-warming, so pre-warm, shared zones
and region resets all target the world the tests run in. Region files are
reset, and chunks force-loaded, only in a
voxelbench_*world; anywhere else that step is skipped with a message and the run goes on. Auto-bench multi-runs follow the same rule. - The start-of-run sweep only removes entities in
voxelbench_*worlds. Weather and time are still set in every world during a run, then restored. /bench zonesshows the zones of the last run, in that run's world, and says so when that world is gone;cleanrefuses zones outside avoxelbench_*world./bench stop, a shutdown and crash recovery sweep onlyvoxelbench_*worlds, with no fallback on the main world.
Tests still remove what they built themselves, wherever they were allowed to run. Two consequences you may notice:
- A server whose worlds hold many loaded entities keeps them during the run, as in normal play.
- A multi-run in a world that is not a
voxelbench_*world — the main world pinned, or Folia without a pinnedvoxelbench_*world — no longer resets regions between runs: it says so, and later runs reuse the chunks already generated there.
Measurements
Folia: three tests measured where they work
On Folia, a test's headline TPS and MSPT come from the regions it works in.
lightingUpdateandtickingTileEntitybuild at their own coordinates, but were measured on the run's shared zones — where they build nothing — as soon as they used more than one zone, the standard benchmark included.playerWorldLoadwas measured at the world centre only, while its simulated players spread up to 8,000 blocks away. Those figures described regions at rest, at about 20 TPS, whatever the load.They are now measured in the regions each test builds in, and
lightingUpdate's Folia measurement starts once its chunks are loaded, as on Paper. Folia results for these three tests may come out lower than in 2.0.2: they now reflect the load the test puts on the server.lightingUpdateandtickingTileEntitycount in the site's Gameplay score, so a standard Folia score can move with them. Paper and Spigot results are unchanged.Custom profiles: loads counted on the zones actually used
In a custom profile, a step whose
dispersedZonesis above 1 always runs on the run's 8 zones. Four tests computed their totals on the step's value instead, so a step at 2 to 7 was measured on the wrong basis:- hopper (Paper and Spigot) filled only the zones the value named — 2 of
8 in
free-host, 4 of 8 inlow-memoryandexample— and the other structures ran empty; a value of 9 or 10 made the test fail. - explosion counted
tntCount× the value as requested, while placing charges in all 8 zones:free-hostplaced 160 charges for 40 "requested". The detonation ratio could reach 4, and the share of charges counted as presented could fall to 0, which lowered the site's confidence in the result. - mobSpawn could publish a survival rate above 100 %. No bundled profile was affected.
- chunkLoading divided its memory cap (30 % of the heap) by the value rather than by the zones it walks: where that cap applied, a step at 2 could load up to four times the chunks the cap allowed.
These totals are now counted on the zones each test actually uses. The standard benchmark, stress-limit runs, single
/bench testruns and steps at 1 or 8 are unchanged. Results of the affected profiles show a step in their history: compare 2.0.3 runs with 2.0.3 runs.Stress profiles:
plateauBoostnow appliesstress.ramp.plateauBoostwas read and published in reports, but never applied: the anti-plateau boost always ran. It now does what the profile says. Only profiles that set it tofalsechange — among the bundled ones,stress-freehost, which now ramps without the boost, as intended for small hosts./bench stresslimitand the other bundled stress profiles keep their ramp.Bundled profiles say what they do
Six bundled profiles (
example,free-host,low-memory,showcase,stress-freehost,stress-monster) now describe what they really do: 8 zones for any step above 1, exceptlightingUpdateandtickingTileEntity, which build exactly as many zones as the step says; theconfig.ymllimits that silently cap chunk loading, mob spawn and explosion steps; six stress types of seven instress-monster;free-hostin English. A copy you never edited is updated automatically; an edited copy is left as it is.A profile's description is part of its fingerprint, the profile hash that
/bench custom infoshows and that voxelbench.com uses to group runs of the same profile.free-host,stress-freehostandstress-monsterhave a new description, so their results from 2.0.3 on are no longer grouped with earlier ones. The other three changed only in comments and keep their fingerprint.Permissions
- A
kind: stresslimitprofile now requiresvoxelbench.stresslimit, like/bench stresslimit;voxelbench.customalone was enough to run one. Stress profiles are offered — in tab completion, on the profile screen and in the list of available profiles — only to those who can run them. - The profile screen requires
voxelbench.customorvoxelbench.start(voxelbench.guialone opened it), and its refresh button runs/bench custom reload, which requiresvoxelbench.reload.
Smaller fixes
- World messages that work when you follow them. A leftover world folder
no longer points you to
/bench world delete, which cannot remove a world that is not loaded; the pinning advice namesvoxelbench_<name>, the name/bench world creategives; the Multiverse message says the world was handed to/mv importand asks you to check/mv list. /bench world listno longer has a[MV]marker: it never appeared./bench verifytells apart a code made for another server, an expired code and too many attempts, instead of "server error, try again later"; a server that is already verified is told so.- Remote monitoring over the plan's server limit now pauses, checks again every 10 minutes and says to turn monitoring off for another server or change plan. It used to retry every minute with a bare HTTP 403.
- A second
/bench linkwhile one is waiting for confirmation is refused. It used to open a second link in parallel, one token overwriting the other. /api/metrics:ram_usedandram_freeare percentages of the maximum heap, and theirunitnow says%instead ofMB. Values are unchanged./bench monitor web startandstatusshow the real address —httpsandmonitor.https.portwhen HTTPS is on.- Chunks VoxelBench force-loads are now released on
/bench stop, on shutdown and after a crash, in every world. A chunk forced when the server crashed used to stay force-loaded for good. - Score screens, chat and the DiscordSRV message use the site's weights — Single-Core 40 %, Hardware 20 %, Gameplay 40 % (they said 45/20/35). They name Redstone and Block Physics under Single-Core, as the site does, and mark the network and single-core CPU tests as not counted in the site's score. VoxelBench shows the site's score; it computes none.
- The
config.ymlcomment on how long voxelbench.com keeps reports now matches the site (new installs only).
Upgrade
Drop the new jar in
plugins/and restart. No config or data migration. Bundled profiles you never edited are refreshed; edited ones are left alone. A group that hasvoxelbench.customwithoutvoxelbench.stresslimitcan no longer run stress profiles: grant the node if it should. A tool reading theunitofram_usedorram_freein/api/metricsnow gets%.- A multi-run could delete regions of your main world.
Fast machines get their score again
A patch release with a single fix: on a fast machine, a benchmark could come back without any score. Everything else is identical to 2.0.1.
What was wrong
Every test sets aside its first seconds of TPS measurement: one second for the monitor to settle, then five seconds of warm-up. The warm-up exists for the gameplay tests, which build their zones and spawn thousands of entities before the real measurement starts — leaving that out is part of what keeps results steady from one run to the next.
The hardware, CPU and network tests prepare nothing in the world, yet they waited out the same six seconds. And the multi-core test does a fixed amount of work, so it ends sooner on a faster machine: on an 18-core server it finished in 6.0 seconds, before the first TPS sample was taken. The test was reported as unmeasured, and no score was computed for the whole report.
What changed
Tests whose load does not run in the world — disk, memory, multi-core, network — now measure from their second second. Gameplay tests keep their warm-up, and the single-core and world-save tests were already timing their own start.
Scores are unchanged. Throughput results such as operations per second or MB/s do not depend on that window; only the TPS figures reported alongside those tests now include seconds 2 to 5, with the main thread idle.
Upgrade
Drop the new jar in
plugins/and restart. No config or data migration.- 2.0.1stablearchivedSep 26, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1, 26.2 · 2.3 MB · 0 ↓No longer available
The disk test no longer pins your CPU or piles up memory
A patch release with a single fix, for the first test of
/bench start: the disk test. Everything else is identical to 2.0.0.What was wrong
The random 4K phases of the disk test allocated a new off-heap buffer for every single read and write — 500,000 of them per phase, two phases per pass, three passes: about 24 GB of native memory reserved over a standard run, for a 1 GB test file. That memory only comes back when the garbage collector reclaims the buffers, and they weigh so little on the heap that a large, mostly empty heap may not be collected for a long time.
- CPU at full load. Allocating and clearing all that memory cost 6 to 7 seconds of CPU per 250,000 operations in our measurement.
- Memory that kept growing. Up to the size of your heap could pile up in native memory during the test — on a container with a tight memory limit, enough to get the server killed.
- A test that never finished. With ZGC and
-XX:+DisableExplicitGC, once the direct-memory limit was reached, every allocation waited for a collection that never came: the disk test stalled with the CPU at full load.
What changed
Each thread of the disk test now reuses one buffer for all its operations. In the same measurement: 0.2 to 0.3 seconds of CPU instead of 6 to 7, and nothing held in memory.
Random 4K results may rise on fast disks: the allocation used to count in the measured time, and no longer does. The other disk results, and every other test, are measured exactly as in 2.0.0.
Upgrade
Drop the new jar in
plugins/and restart. No config or data migration. - 2.0.0stablearchivedSep 25, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1, 26.2 · 2.3 MB · 2 ↓No longer available
See what lags your server and what fills its memory — on measurements you can trust
VoxelBench 2.0.0 brings together everything since 1.8.0. If you are coming from 1.8.0, everything below is new to you.
It adds two diagnostic tools to the plugin.
/bench profileshows which plugins, which parts of the server and which methods keep your tick threads busy, using Java Flight Recorder — built into Java, nothing to download./bench memoryshows which plugins fill the heap, takes a full heap dump when you need one, and analyses that dump to name the plugin holding the memory. A profile or a memory report can be shared by public link, like spark, or kept in your voxelbench.com account; a heap dump never leaves your server.Remote monitoring is rebuilt: you turn it on from the game, the voxelbench.com dashboard sees each interval as a whole instead of one instant out of sixty, follows real time on a lagging server, and on Folia measures every region on its own. Tick, garbage-collection and CPU metrics are measured by VoxelBench itself from startup, without spark.
And the benchmark was audited against its own claims. Several hardware metrics measured their own instruments more than the hardware, the standard benchmark could be reshaped by the measured server's configuration, most permissions were declared but never checked,
/bench testbuilt and dug in the world you stood in, and on Folia auto-bench did not run at all. All of that is fixed.Scores are not directly comparable with 1.8.0 for memory and disk results, for two gameplay tests, and on servers that lag under load; the standard gameplay load itself is unchanged. Compare 2.0.0 runs with 2.0.0 runs — the details are in the Upgrade section.
A profiler in the plugin
/bench profile [seconds]samples the tick threads — the main thread, or every region thread on Folia — and prints, as percentages, which plugins and which parts of the server they spend their time in, and the busiest methods. Self counts the code on top of the stack, including the server code a plugin calls; inclusive counts anything below it. Libraries a plugin declares inlibraries:and plugins declared withpaper-plugin.ymlcount for that plugin.- A ring buffer that catches lag by itself.
/bench profile ring onkeeps the last minutes of samples. When the server lags — 40 ticks in a row of 100 ms or more, or a single tick of one second, by default — VoxelBench dumps and analyses that window on its own, at most once every five minutes. You read the profile of a lag spike that already happened. - In the menus and on the web dashboard. A Profiler button in the main menu starts a capture (30, 60 or 120 s, then keep, share or upload it) and controls the ring buffer; the Reports screen gets a Profiles archive — each profile with its plugins, methods and warnings, and buttons to share, upload, remove the link or delete it. The embedded web dashboard shows the same archive, read-only.
- Profiles stay on your server, next to your reports in
plugins/VoxelBench/reports/profiles/, with their own retention limit. - Benchmarks are never profiled. During
/bench start, a stress-limit, tier or auto-bench run, or a test whose result is sent to voxelbench.com, captures are refused and the ring buffer stops; profiling resumes 30 seconds after the run. - Results are sample counts and percentages, never durations: the sampler cannot promise how often it fires, so turning samples into milliseconds would invent precision.
Share it, or keep it in your account. Nothing leaves the server unless you ask, and each command sends one profile right away —
lastnames the newest,previewshows what would leave without sending anything, and adding the verb to a capture sends its result as soon as it is saved (/bench profile 30 share)./bench profile share <id>: a public link, like spark. The profile is published behind an unlisted link that anyone who has it can open, for 7 days, without any account. What leaves is a cleaned copy: IP addresses, host names, URLs, file paths, UUIDs, tokens and the names of online players, worlds, the OS user and the machine are removed, and the chat says what was removed./bench profile unshare <id>deletes the link before it expires. Only official builds can share./bench profile upload <id>: private, to your account. On a server linked to voxelbench.com, the profile is kept in the owner's account, as it is.- Neither ever happens automatically or during a benchmark.
The profiler is marked experimental and needs Java Flight Recorder:
/bench profile statussays when a minimal Java runtime or the JVM options leave it out. Permissionsvoxelbench.profile, and on top of itvoxelbench.profile.shareandvoxelbench.profile.upload, operators by default.Memory inspection
/bench memoryis what spark'sheapsummaryandheapdumpdo, built in, plus an analysis of the dump that spark leaves to external tools. A summary or an analysis stays on your server unless you share or upload it; a heap dump never leaves it./bench memory summarycounts the objects of every class on the heap and attributes them to each plugin, the server or Java. It freezes the server about a second for 1.5 GB of heap in use, two for 3.5 GB (half that withall, which skips the garbage collection but counts garbage too), and saves each summary next to your reports, inreports/memory/./bench memory dumpwrites a full heap dump (.hprof) to open in a memory analyser. It freezes the whole server while it writes — from seconds to minutes depending on heap size and disk speed — so it first shows the expected freeze, the file size and the free disk space, and runs only once you confirm. Addgzipto compress it afterwards, with the server running.- A dump that would get the server stopped is refused. Past the watchdog
limit (
settings.timeout-timeinspigot.yml, 60 s by default), the server considers itself crashed and stops. When the expected freeze reaches that limit, the dump is refused with the reason and the way out: raise the limit, or addforceif the server may stop — the preview then repeats the warning in red, andforcelifts nothing else. - A heap dump holds everything in memory — tokens, plugin passwords,
players' data. It is written to
plugins/VoxelBench/heapdumps/, outside the reports folder, readable only by the server's account; at most two are kept, and VoxelBench has no command, button or automatic path to send one. /bench memory analyzereads the dump for you — or addanalyzeto the dump command to do both in one go. A summary counts a plugin's cachedbyte[]and strings for Java; the analysis follows the references and shows the memory each plugin retains, the biggest leak suspects and the chain of fields that keeps each one alive. In our tests, a plugin caching 500 MB showed 0 % in a summary and about half the heap in the analysis, found in 15 to 40 seconds for a 1.5 GB dump.- The analysis never runs inside the server. It runs in a separate Java process at the lowest priority, with one CPU core; TPS stayed at 20 during our tests on a 12-thread machine (on 2 CPU threads or fewer, where it competes with the server, the chat warns first). Before it starts, VoxelBench checks the free memory and the container's memory limit, so the analysis cannot get your server killed for lack of memory: when the full analysis does not fit, a quick one runs instead (sizes per plugin, worlds, chunks and entities in memory, duplicate strings), and when not even that fits, it is refused with the figures. The report holds class, field and plugin names and sizes only, never a value from the heap.
- Big heaps get a full analysis too. The memory of the full analysis follows
the number of objects (about 90 bytes each), not the file. When it does not
fit, the analysis keeps its object graph on disk, in work files it removes
afterwards: about half the memory for exactly the same report, two to three
times slower (60 s instead of 26 s for a 2.2 GB dump of 36 million objects in
our tests), and down to a quarter of the memory (0.8 GB, 2.5 minutes) as a last
resort.
memory-inspection.analysis.disk: falseturns it off. - What one plugin retains. When you suspect one plugin,
/bench memory analyze last retained <plugin>answers how much it retains — what would be freed without it — and the heaviest classes it holds, with a fraction of the memory of a full analysis. - Fewer false alarms. The static state of the server's own classes, which retains 10 % of a small heap on its own, is shown apart instead of as a leak suspect — recognised from the shape of the heap, so Spigot, Paper, Folia and their forks behave alike — while a single class whose static cache keeps growing is still a suspect. Plugins are recognised even when a fork or a hybrid server loads them with its own class loader. And in a full analysis, duplicate strings count only the copies still in use.
- Safe on hosting panels. The analysis shares the container's memory, CPU
and disk limits with your server. When the server's heap leaves no room (the
default Pterodactyl setup gives it 95 % of the container), the chat says which
-Xmxwould. Setmemory-inspection.host-disk-limit-gbto your plan's disk limit and dumps and work files count against it — a panel stops a server that goes over, and VoxelBench warns when it sees one without the setting. With a CPU quota of 2 cores or less, the analysis takes only 30 % of one core, so the quota stays with your server. - Share a report, or keep it in your account — like a profile.
/bench memory share <id>publishes a cleaned copy of a summary or an analysis behind an unlisted public link, for 7 days, without any account: plugin, class and field names, sizes and leak suspects stay, while addresses, paths, UUIDs, script names, the names of players, worlds, the OS user and the machine, the time zone and the machine's memory figures are removed./bench memory upload <id>keeps the report, as it is, in the account of a linked server.last,previewand the verb added to the command that makes the report work as for profiles —/bench memory summary live share,/bench memory analyze last upload,/bench memory dump live analyze share— and for a dump, the analysis leaves, never the dump./bench memory unshare <id>deletes the link. Never automatic, never during a benchmark; only official builds can share. - In the menus: a Memory button next to the Profiler takes a summary or analyses the last dump, and a Memory tile in the Reports screen lists the saved summaries and analyses and shows, for each, which plugins hold the memory, with buttons to preview, share, upload or remove the link. Heap dumps stay command-only.
- None of this runs during a benchmark. HotSpot JVMs only (OpenJ9 has neither
the histogram nor the dump). Permissions
voxelbench.memory, and on top of itvoxelbench.memory.dumpfor dumps and their analysis,voxelbench.memory.shareandvoxelbench.memory.uploadto send a report, operators by default. VoxelBench's own footprint went down too: it now keeps its language files in about 1 MB of heap instead of 6.
Forms, and confirmations without retyping
- Forms instead of options. Typed alone,
/bench profile,/bench memory summary,dumpandanalyzeopen a form — duration, interval, live or all objects, gzip, analysis mode, which dump, keep, share or upload — that builds the command for you: a native dialog on Paper 1.21.6 and later, an inventory screen elsewhere. Only what your permissions andconfig.ymlallow is offered, and a dump still shows its preview and asks for confirmation. - Each form reopens on your last choices, even after a restart. Keeping, sharing or uploading the result, forcing a dump and which dump to analyse always start from their default: nothing leaves because you chose it last time.
/bench confirmconfirms your last preview — a dump, or a profile or reportpreview— from the console too, with all the checks of the full command. Players also get a Yes/No window with the preview (confirmation.popup: falseturns it off).
Remote monitoring: on from the game, the whole interval, in real time
Turning on remote monitoring — the live dashboard on voxelbench.com — meant editing
config.ymland restarting; enabling it before the site was ready switched it off for good, without a word in game. And what it sent was one instant of each minute, counted in ticks./bench monitor remote on|off|status(permissionvoxelbench.monitor.web).onlists what will be sent, checks in with voxelbench.com right away and tells you what is left to do: nothing, enable monitoring for this server on the site, check your plan, or link the server again.statusshows the link, whether monitoring is sending, paused (and why) or stopped, and what the site kept or refused.- Any order works. While the site is not ready, monitoring pauses and checks
again on its own — every minute, or every 10 minutes when your plan does not
include it.
/bench reloadapplies its settings, and the plugin never rewritesconfig.ymlitself. - Each snapshot describes its interval, not one instant. Each snapshot now carries the interval's TPS, the average, 95th-percentile and worst tick time, the ticks over the 50 ms budget and the seconds spent below 18 TPS. A frozen server reports 0 TPS; a server not measured yet reports nothing instead of a made-up 20.
- Real time, even when the server lags. Snapshots, uploads and the heartbeat were counted in ticks: at 10 TPS a one-minute snapshot became a two-minute one, alert rules went quiet exactly when TPS collapsed, and a lagging server could show as offline. They now follow the wall clock, and uploads happen every minute by default.
- Folia: every region, on its own. On Folia the TPS came from the global
region, which is almost idle. Every region Folia ticks is now measured — farms
and chunk loaders without players included — and listed worst first with its
world, centre block (as Folia's
/tpsshows it), chunks, players, entities, TPS and tick times. The headline figures carry the worst region; averages across regions come alongside. - A frozen server is visible, a sleeping one is not an alarm. The heartbeat
says how long ago the server last ticked. Since Minecraft 1.21.2 an empty
server pauses after
pause-when-empty-seconds: the heartbeat says so instead of looking frozen, and never while a player is online. - Memory that explains a crash: heap still used after garbage collection, the memory the process really uses, and the container's memory limit — a process close to it is killed by the system without any Java error.
- Nothing lost on the way. A full buffer no longer drops a snapshot, a failed upload is retried in order, and the last minutes before a shutdown are sent before the plugin stops. Turning monitoring off is announced, so the dashboard does not raise an "offline" alert, and an event category your plan does not include is silenced alone instead of being refused on every event.
- The
remote-monitoring.eventssettings, which had no effect at all, now apply, and the GC major event counts only real major-collection pauses.
Tick, GC and CPU metrics without spark
VoxelBench now measures every tick from startup, whether or not spark is installed: tick durations from the server's own tick event on Paper and its forks, from the main thread's CPU time on Spigot. Garbage collections are read one by one — major or minor, longest pause, collector and cause — with an estimated allocation rate.
- The embedded dashboard's
/api/metricsgains aperformancesection with 10-second, 1-minute, 5-minute and 15-minute windows. - New PlaceholderAPI placeholders:
%voxelbench_mspt_p50%,_p95,_p99,_peak(with an optional_10s,_1m,_5mor_15msuffix),%voxelbench_tps_10s%to%voxelbench_tps_15m%, and%voxelbench_tick_source%.
Measurements you can trust
Memory latency measured the clock. The loop surrounded a single byte load with two clock reads, ten million times over. Reading the clock alone costs about 41 ns; the published figure was 48 ns — roughly 7 ns of memory signal. It also could not tell cache from RAM: the same code returned 47, 46 and 42 ns on a 256 KB, an 8 MB and a 512 MB buffer. The new measurement walks a chain where each load depends on the previous one: on the same three buffers it returns 14, 105 and 199 ns — what a memory hierarchy actually looks like.
- Random 4K disk I/O measured the page cache. Nothing bypassed the operating system's cache, and the reads targeted the blocks written seconds before. The 4K phases now use aligned direct I/O, and are published only when the cache bypass actually succeeded.
- Four timed loops were timing a random number generator. Memory sequential write spent 69 % of its time generating data; disk sequential write could never exceed the generator's own 290 MB/s, whatever the disk. Data is now prepared before the clock starts. Bandwidth reports the best pass instead of the average, and the spread between passes is published so a noisy machine and a slow one no longer look alike.
- These metrics carry new names in reports (
memoryAccessLatencyNs,directRead4K_*,directWrite4K_*), so an old value and a new one can never be ranked side by side. - The lag filter discarded the lag. Every sample below 15 TPS was thrown away as "not representative". A server at 8 TPS is a server at 8 TPS; sustained lag is now measured, while genuine artefacts — catch-up ticks, GC pauses, a drifting sampler — are still excluded.
- The warm-up never protected a single measurement: it was reset right after
being set.
lightingUpdatemeasured its own preparation instead of the lighting work;tickingTileEntitylet part of its load stop ticking. - The standard benchmark measures the same thing everywhere. Three tests of
/bench startreadconfig.ymlwhile building their load, and so did the number of zones: an operator could reduce the load of a "standard" score without the report saying so. It now always runs the load every default installation has been measuring — 10 hopper lines of 50, 150 TNT, 2,000 chunks per zone, 8 zones — so default installations see no score change. The limits still apply to/bench test, custom mode and custom profiles. - Reports record the server's difficulty: Peaceful removes hostile mobs from the entity tests, so a comparison can see when two servers were not measuring the same thing.
/bench teststays out of the world you stand inMost single tests ran in the launching player's world — usually the main world.
lightingUpdatedug shafts,chunkTickingcleared about 545 × 545 blocks,explosion,hopper,redstoneand others built and cleared their zones there,mobSpawnremoved every mob within 100 blocks of its zones,chunkLoadingleft thousands of generated chunks in your world files, andplayerWorldLoadleft a hole for every block it touched near spawn.- Every test that writes to the world now runs in a dedicated world: the one
pinned with
/bench world set, otherwise a temporary flat world deleted at the end — even after a failure or/bench stop. When neither is possible, the test is refused instead of falling back to your world. - You are brought back where you were. Tests that need a player nearby take the player who launched them to their zone and bring them back to their position and game mode when the test ends, fails or is stopped — and someone who disconnects mid-test reconnects where they were.
playerWorldLoadrestores every block it changes to its original state, and never touches chests, spawners, signs or beds./bench start, custom profiles, stress-limit and tier runs are unchanged and skip no test.
Custom profiles run the load they ask for
The hopper limit meant for
/bench test hopper <lines>silently capped every profile: the bundled profiles asked for 20 to 100 lines and all ran 10. Profiles now run what they ask, up to 500 lines per zone, and the bundled ones say 10 — the load they have always run — so none of their results move.- New bundled profile
hopper-heavy: 100 lines in each of 8 zones, ten times the standard load, for servers with large sorting systems. Heavy, and not comparable with/bench start. - Bundled profiles reach existing installations. A new bundled profile is
copied once; one you delete stays deleted; an untouched copy of an older
version is updated — including the
showcaseprofile that 1.8 always rejected — and a copy you edited is never overwritten.
Folia
Folia — and the forks built on it, such as Canvas — splits the world into regions, each ticked by its own thread, and refuses outright an operation on an entity from any other thread. Seven operations got that wrong.
- Auto-bench ran again.
/bench autobotvalidated its job over HTTP, then resumed on the global region, where putting the player into spectator threw on the spot: on Folia, auto-bench did not run at all. - Nobody is stranded in a benchmark world. A player reconnecting inside a leftover benchmark world is teleported out; that teleport failed on Folia before restoring the game mode, leaving the player in spectator with nothing able to get them out.
- Stress-limit runs restore the player's position and game mode,
/bench zones tpworks, startup cleanup of a previous run's leftovers actually runs, the scoreboard is reassigned when a player reconnects mid-run, and/bench zones cleansweeps each region on its own thread. - Gameplay tests that write to the world need a world pinned with
/bench world set, since Folia cannot create worlds while running; without one they are refused with a message that says how to pin one.
Permissions that hold, settings that apply
Most permissions in
plugin.ymlwere declared but never checked: anyone withvoxelbench.usecould start a benchmark or a stress-limit run, stop someone else's run, or open the monitoring and report screens. They are now enforced on the commands, in tab completion and on every menu screen.- Every built-in test has its own node, and starting a test by its full name no
longer skips it.
/bench linkrequiresvoxelbench.link; teleporting to and clearing test zones requirevoxelbench.world; the settings screen requiresvoxelbench.settings. config.ymlcontained thebenchmark:section twice, and the second silently replaced the first; the file is repaired on upgrade, with a warning.dispersed-zones.defaultnow applies to custom mode; settings that nothing read were removed, and settings the plugin used were added with their defaults. New settings are always added to your file; nothing you set is overwritten.
Security, reports and integrations
- The embedded web dashboard is private by default. It listens on the server
machine only (
monitor.bind-address: "127.0.0.1"). Listening on any other address requires a dashboard password, set from the server console only — a command typed in game ends up inlatest.log. Passwords are stored with PBKDF2, and the login page follows the IP whitelist and turns an address away after five failed attempts in a minute. - LiteBans: only real sanctions. The hook read the commands players typed,
before any permission check: anyone typing
/ban <name> <text>produced a "player banned" alert, even when LiteBans refused the command. VoxelBench now listens to LiteBans' own events. On new installations LiteBans forwarding is off by default, and the typed reason is sent only if you enableremote-monitoring.events.litebans.include-reason. - Minecraft 1.17 submits reports again. With the default anonymization setting, preparing a report called a JSON method that 1.17 does not ship, so submission failed — and server verification, update checks and remote monitoring with it. The plugin is now built against the exact library version 1.17 ships.
- A benchmark is saved locally at the end of every run, whatever happens to the submission, with its status and the reason when it failed. With full anonymization, file paths inside JVM flags are masked too.
- No more invented numbers. Sub-scores the service never computes no longer
show as
0(the placeholders returnN/A), and scores without an upper bound are no longer drawn as bars out of 100.
Smaller fixes
Auto-bench: stopping a stress-limit or single-test job works, and jobs that cannot start fail at once with the real reason instead of waiting for a timeout. Folia: chunk preloading no longer defers every load by a tick, and Folia 1.20.x no longer logs a stack trace at every monitoring collection. Remote-monitoring requests go straight to the right address. Recovering after a crash no longer sweeps the main world when the run's temporary world is gone. The whole interface is translated — monitoring, custom profiles and report screens included — and a saved report reads in the viewer's language.
Upgrade
Drop the new jar in
plugins/and restart.config.ymlis repaired and migrated on first start; new settings are added, and nothing you set is overwritten.- Comparability (against 1.8.0): memory and disk results change scale and
names,
lightingUpdateandtickingTileEntitynow measure the load they announce, and servers that lag under load report that lag. The standard gameplay load is unchanged. Single gameplay tests launched with/bench testnow run in a flat benchmark world: they match what/bench startmeasures, but not your earlier/bench testruns. - Permissions: operators and holders of
voxelbench.*are unaffected. A group that only hadvoxelbench.useloses the guarded commands until it is given the corresponding nodes.voxelbench.test.mobpathfinding,.redstoneand.blockphysicsmoved from the single-core group to the gameplay group, their actual category. - Configuration: if a warning mentions the duplicated
benchmark:section, set again any value you had placed in its first copy. Monitoring events now use the thresholds yourconfig.ymlalways displayed: TPS drop after 10 s, TPS critical at 10.0, GC major at 200 ms of major pauses. - Web dashboard: if your
monitor.bind-addressis0.0.0.0(the former default) and no dashboard password is set, the dashboard no longer starts. Set a password from the console (/bench monitor auth password <password>) or bind it to127.0.0.1. - Remote monitoring: if an earlier version switched it off after the site
refused it, turn it back on with
/bench monitor remote on. Asend-intervalstill at the former default of 300 seconds becomes 60; acollect-intervalbelow 30 seconds is raised to 30. - LiteBans: an existing installation keeps forwarding sanctions if it did, but
without the typed reason until
remote-monitoring.events.litebans.include-reasonistrue. - Custom profiles: a profile of yours whose hopper step asks for more than the command limit now runs that many lines — its earlier results used fewer. VoxelBench names each such profile in the console at startup.
/bench test: a gameplay test launched from your main world now creates a temporary world for itself. To avoid that, or on Folia, pin a world loaded at startup with/bench world set— the test builds and clears its zones there. Auto-bench jobs that run a single gameplay test follow the same rule.- Profiler and memory inspection: operators can use them right away; the ring
buffer stays off until you turn it on. Public links and uploads can be turned
off with
profiling.share.enabled/profiling.upload.enabledandmemory-inspection.share.enabled/upload.enabled;memory-inspection.analysis.max-heap-mb(4096) andmemory-margin-mb(512) bound the memory an analysis may take. On a hosting panel, setmemory-inspection.host-disk-limit-gbto your plan's disk limit. If your backups or shared folders includeplugins/VoxelBench/, leave outheapdumps/.
Java 16 or newer for Minecraft 1.17 to 1.21.x; Java 25 for Minecraft 26.1 and later. Tested on Spigot, Paper and Folia, up to Minecraft 26.1.2.
- 1.8.0stablearchivedSep 6, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1, 26.2 · 1.6 MB · 20 ↓No longer available
This release fixes one defect that ran through the whole bench, though its symptoms looked nothing alike: runs that stopped without a message, tests cut short on slow hosts, and entire reports thrown away while most of their measurements were perfectly good.
Scores remain comparable with 1.7.x. No formula changed, and a report produced by 1.7.0 still compares directly to one produced by 1.8.0.
A deadline counted in the unit the failure destroys
Between two tests, VoxelBench waits eight seconds for the machine to settle. That delay was expressed in ticks — the server's internal clock. A tick lasts 50 ms when all is well, and a great deal longer when the server is struggling.
So the window stretched exactly in proportion to how useless it had become: the previous test had just saturated the machine, and that was precisely why we were waiting.
Measured on a Folia server during the mob-AI test, at 0.26 ticks per second — one tick every 3.8 seconds — the old window's 160 ticks would have taken 615 seconds: more than ten minutes for an eight-second pause. It now takes 8064 milliseconds.
The change only touches waits whose subject does not depend on the server's rhythm. TNT fuses, redstone timing, fluid propagation and per-tick chunk budgets stay in ticks: that is their correct unit, and converting them would have distorted the measurements instead of protecting them.
Runs no longer stall between two tests
A reported log showed a benchmark stopping dead after the chunk-loading test. No error, no message — and the global lock held until the server was restarted.
The cause: nothing watched the gap between two tests. One test's watchdog is disarmed the moment it finishes; the next one's is only armed when it starts. In between sit the teardown of thousands of chunks and the entire setup of the test to come.
A watchdog now covers the run as a whole, and its thresholds are derived from the limits each test declares — a hand-picked threshold would have cut the memory test short, which legitimately asks for up to fifteen minutes.
A test that proves it is progressing earns more time
The memory test had a fixed 300-second limit. On a slow host its final pass needed 377: the test was cut, and the whole report refused.
A test can now request a reprieve. Each completed unit of work pushes the deadline back, while an absolute ceiling stays in place — the point is not to remove the protection, but to tell a slow test from a stuck one.
The reprieve renews on evidence: a pass that genuinely finished. Never on a mere sign of life, or the watchdog would push back its own deadline forever and stop being worth anything.
Messages now say what happened, and what to do
A user read this in their chat:
⊘ Test ignoré: Mémoire — memoryPressure: projected peak heap 107.6% after pre-flight GCA developer's sentence, in English, in the middle of a translated message. It names a cause and never a remedy.
Every test that does not run to completion now carries a structured reason: a translated cause, and above all a computed remedy. No longer "raise
-Xmx", which leaves you guessing, but the amount of memory you would have needed, worked out:⊘ Test skipped: Memory — not enough memory — projected peak would reach 107.6% of the heap (1670 MB available) → Raise -Xmx (about 2304 MB for the standard profile) or run /bench start low-memory.Two additions complete this. A warning before the run when memory plainly will not be enough: the verdict was computable in the first second, and the user learned it in the eleventh minute. And an end-of-run summary that ties the skipped tests together — three notices six minutes apart do not connect themselves in a scrolling chat.
A partial run is no longer lost
This is the most visible change for anyone measuring a small server.
When memory runs short, VoxelBench refuses to start a test rather than bring the server down. That is the right decision. But until now, a report missing a required test was rejected outright: a host with 1670 MB of heap lost twelve correctly measured tests because three had been skipped — that is, because the plugin had correctly protected itself.
The report now carries, for each test, a status and a cause code. The service can therefore tell a deliberate skip from a crash, and the report is kept and viewable.
It is not scored, and that will not change: putting a number on a run with a missing required test would mean scoring something nobody measured. It does not appear in the leaderboards — but you no longer lose eleven minutes of measurement.
What the screen claimed without any way of being wrong
The
/bench test diskcommand reported "✓ Test passed", "TPS 20.00" and "MSPT 50.00 ms". Those three values were hard-coded: they appeared regardless of the actual result.Worse, all six throughput metrics displayed 0 MB/s on an NVMe drive that had just measured 2621 MB/s sequential read. The measurement was correct — it reached the report intact — only the display lied, and in the costliest direction for someone whose whole reason for running it was to diagnose their disk.
A missing metric now shows "—" rather than "0". Zero is a plausible throughput: confusing the two reads as a dead disk where there is only a missing value.
Quieter fixes
- A red error line (
No key layers in MapLike[{}]) was written to the console on every run, by VoxelBench's own bench-world creation. The fix existed but had only been applied to Folia; three Paper logs said otherwise. The generated terrain itself was always correct. - The memory collection performed before each test abandoned its loop on the first attempt, mistaking "there is nothing left to free" for "collection has not started yet". It now waits for proof that a collection happened, and says so when that proof never arrives.
- The memory-leak detector reported leaks that did not exist: it compared a test's end to its start, and so measured what the test had just created and not yet cleaned up. A leak is what survives cleanup.
- A test now waits for the previous one's chunks to actually unload, rather than for a delay to elapse. When the wait reaches its limit, the log says so instead of starting silently on a cluttered world.
- The pre-flight check logs that it is waiting. When it opened its confirmation screen, the command stopped without leaving a trace: the server answered "no test running" and nobody could tell why.
Compatibility
Validated on three platforms, full run and report submitted: Paper 1.21.11, Spigot 26.1.2 and Canvas 26.1.2 (Folia family). Continuous integration additionally covers Spigot from 1.17.1 to 26.2.
The new report fields are optional: a service that does not know about them ignores them, and reports produced by earlier versions are read exactly as before.
- A red error line (
- 1.7.1stablearchivedAug 17, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1, 26.2 · 1.6 MB · 38 ↓No longer available
This release fixes what the bench was measuring, not how it scores. Three defects compounded, and together they meant the TNT stress scale never broke on any machine: it climbed to its configured ceiling on every host, measuring entity spawning and platform construction rather than your server. The headline is short — the bench now proves that a test ran before it publishes that test's number.
The explosion test detonated nothing: forcing a chunk loads it but never brings it to
ENTITY_TICKING, so fuses sat at 40/40 after ten seconds while the charges were reported "alive". Underneath, the measurement layer failed open — sixteen fallbacks across the region monitor and the tier verdict all meant "healthy server", so a tier that measured nothing was published as20.00 TPS / 0.00 ms, the best possible result. And under that,max-tnt-per-tick: 100, the default on Spigot and Paper, silently capped how much load ever reached the server.protocolVersionstays 1.2.0 andAPI_VERSIONstays 1.2.0: the wire contract is additive. What changed is what the numbers mean.Your previous results do not carry over
Benchmark and stress-limit results from 1.6.0 and earlier are not shifted. For
stressTnt(every platform) andstressHoppers(Paper, Spigot) they are void, and no Folia stress-limit result existed at all before this version. There is nothing to convert and nothing to re-read: no 1.6.0 report carries a single validity witness, so it is impossible to tell after the fact which of these causes applied to it.Re-run your campaigns, and expect lower limits. That is the measurement becoming real, not your machine regressing. Scale bounds moved as well, so a tier number and a load value no longer mean what they meant.
Measured breaking points on the corrected TNT scale, for reference: 1938 ±76 (Folia), 2229 ±92 (Paper), 2279 ±65 (Spigot). The corrected hopper scale breaks at 3112 ±71 on Paper, where it previously returned 40 000 — against 121 on Spigot and 2 661 on Folia, three mutually incompatible answers for one test.
Reports produced by 1.7.0 pre-releases are not comparable to this build either. Reject any that show
entityTickRatioof exactly1.00on a test with transient entities, orworkCompletedof exactly0; both were wrong by construction until late in the cycle.The explosion test now actually explodes
Forcing a chunk attests to loading, not to ticking, and only a ticking entity burns a fuse. Charges were created, accepted by the server, then frozen — and since frozen work costs a region nothing, TPS held at 20.00 while the scale climbed. Zones now hold their own ticking lease, released zone by zone (one zone finishing used to release the force-load of the other three, which stopped ticking mid-flight), and the measurement window does not open until the zones really tick. At tier 0, 498 blocks now fall where zero fell before.
On Paper that promotion takes about 20 seconds — measured, 9 → 200/200 chunks in 17.8 s on a 150×8 run, against 1.8 ms on Spigot. The wait is announced in console every 5 s, on the scoreboard and in the action bar, capped at 30 s, and counted outside the measurement window: it does not enter
durationSeconds,blocksPerSecondor the TPS/MSPT window. It is not a freeze.The gap between "loaded" and "ticking entities" is now measured rather than assumed. Over 8 zones and 200 forced chunks, Paper returned 88
ENTITY_TICKINGagainst 112TICKING— a level that runs blocks and random ticks, so it has every appearance of an active zone — while Spigot returned 200/200.A tier that measured nothing is no longer the best tier
The "no result" branch emitted 20.00 TPS and 0.00 ms across fourteen fields with
measurementDegraded: false: the most degraded case of all passed for the only healthy one. Symmetrically,loadAbsorbedRatio: -1— meaning "no witness" — was remapped to 1.0, so the tier was not saturated and it passed, word for word what that method's own documentation said it existed to prevent.The verdict now distinguishes "no witness" from "a witness saying we cannot judge" and returns
NOT_MEASURED— neither held nor broken — read before the health branch. A tier whose load was never censused therefore stops being labelled "health score low" in a message stored in the database and shown to you, as if your machine were at fault.Two consequences follow:
- A tier the server could not detonate within its window is now a broken tier. Explosion saturation was invisible to every other metric: queued work costs a region nothing, so TPS held at 20.00 and the degraded-second fraction at 0.00 across thirty-two tiers while four charges in five were never touched (798 detonated + 3202 still alive = 4000 requested). The tier now breaks below 95 % presented load, and the break reason names it.
- A collapsed region publishes its own collapse. The witness sampler is
scheduled in ticks of the measured region, so a frozen region never runs it and
the measurement vanishes exactly where it matters. A scheduled task that never
started in D seconds proves the region executed fewer than 20 ticks over that
span, which bounds its rate from above with no sampler at all. Published as
regionTaskStarvedandregionTickCeiling— on mobAI, under 0.68 tps, corroborating the 0.30 tps the region monitor measured independently.
Your server's own TNT cap no longer caps the run in silence
max-tnt-per-tick: 100is the default on Spigot and Paper. A test posting 1200 charges never saw most of them tick, so it measured a configuration file rather than your machine — silently, on almost every host.The cap is now raised in memory for the duration of the test and restored on every exit path, including
/bench stop, a timeout and an exception. Yourspigot.ymlis never written, and a restart re-reads it intact. Measured on 150×8 over two consecutive runs per platform, the second run proving the first restored:tntNeverTicked300 → 0,tntPresented900 → 1200/1200,detonationRatio0.89 → 1.00,loadAbsorbedRatio-1 → 1.00.Reports carry
hostMaxTntPerTick,hostTntQuotaSourceand one ofhostTntQuotaRaised/hostTntQuotaLimited/hostTntQuotaUnknown, so two runs of the same server stay distinguishable. Setbenchmark-tests.explosion.raise-host-tnt-quota: falseto refuse any in-memory change; the run is then reported as partial and a console warning names the value to raise yourself.The read itself was broken too.
Bukkit.spigot().getClass().getMethod(...)resolves an override declared by a non-public anonymous class, soinvokethrewIllegalAccessExceptionon every call, swallowed by a catch — and every read fell back tonew File("spigot.yml"), relative to the JVM's working directory, missing as soon as a panel or a systemd unit sets another one. There are three ordered routes now, and the config-shape check refuses any file that is not aspigot.yml, because reading a falsely high cap elsewhere would credit a host with absorbing load it never received.Every report now carries the evidence that it measured something
Each test publishes a witness saying whether its entities actually moved, how many of its chunks really ticked entities, and how much of the announced load the server was actually given. New fields include
entityTickRatio,tntPresented,tntNeverTicked,detonationRatio,detonationUnaccounted,explodeForeign,zoneChunksEntityTicking,workCompletedwith its unit,loadAbsorbedRatio,sampleCountand per-tiermeasurementDegraded.systemInfonow also distinguishes the host machine from what the server can actually use, throughavailableProcessors,memoryLimitGbandprimaryDisk.usableGb/.totalFsGb.entityTickRatiois built onticksLived, the one counter that only advances when an entity is ticked. It is armed on every test, sampled over at most 256 entities and collected once a second from each owning region, so a partial freeze stays visible;entityTickSamplesandentityTickConclusivesit beside it, because "the sampler never ran" and "nothing was readable" are different failures. For explosions the fuse is the exact witness — it decrements once per entity tick, so a charge still at its initial value was never ticked once. The estimate it replaces returned 694 charges presented against 800 detonated, a ratio of 1.15 that novalue < thresholdguard could ever catch.Two rules follow for anyone reading a report:
- Read the flags before the numbers.
loadNotPresented,loadPresentationUnmeasured,zoneTickingNotReady,zoneTickingCapped,entitiesNeverTicked,entitiesPartiallyTicked,entityCensusPartial,regionTaskStarvedandmeasurementDegradeddescribe a doubtful measurement, not a slow machine, and that tier does not compare to the others. They are raised only when something is wrong: a healthy report carries none, because a permanent flag is noise and noise stops being read. - Never read
-1as0. It means "cannot conclude", never "nothing moved".
On a populated server, TNT lit by your players or your redstone is now counted separately as
explodeForeigninstead of being credited to the machine under test. The old probe incremented for anyTNTPrimedexplosion anywhere on the server, and that counter decides the tier verdict: on the same host the same day, 1 zone gave 1.00, 8 zones gave 0.67, and 8 zones with simulation distance raised gave 0.48 with no survivors — three limits for one machine, driven by a bench parameter. Still, run your benchmarks on an idle server: foreign detonations no longer distort the verdict, but they contaminate the run.Warm start, and a break that has to repeat
The last stable value of each stress test is remembered in
data/stress-warmstart.yml, and the next run starts near 70 % of it instead of climbing the whole scale again. If the warm probe breaks, the runner refines downwards; if nothing there holds, it falls back to a clean full climb. Stale and other-algorithm entries are ignored. Tunable understress-limit.warm-start:enabled(true),factor(0.70),max-refine-down-steps(6),max-age-days(30).Warm start makes a run depend on the previous run of that machine, so two successive runs of the same server are no longer directly comparable with it on. Set
stress-limit.warm-start.enabled: falsewhen you are comparing two servers against each other.Separately, a broken tier is now replayed before the scale stops, on the way up and during refinement, with the confirmation variance tightened from 30 % to 12 %. An isolated hiccup no longer condemns a healthy machine. Runs take longer, and some limits go up relative to 1.6.0 for this reason.
/bench tier — walk one scale on its own
/bench tier <test> [tiers=N] [from=N] [to=N] [duration=SEC] [zones=N] [cold] [noskip]runs a single stress test up its scale to diagnose it, without launching a full campaign. It sits behind the newvoxelbench.tierpermission (op by default), never writes the warm-start memory, and submits on the unit-test channel governed byreports.backend.unit-tests— never as a stress-limit report. Note thatfromandtoare loads, not tier numbers.Folia
Headline TPS and MSPT now describe the measured window instead of being an exponential-moving-average snapshot read at the worst moment of the run, so every Folia headline number changes magnitude.
An entity the current thread cannot read is also no longer counted as a dead one. Reading an entity from a thread that does not own its region throws under Folia, and the old code turned every missed read into a corpse — which produced "no charges left", then "so they exploded", then "so a plugin is cancelling the damage": three conclusions stacked on missed reads.
EntityCensusnow publishesentitiesTracked,entitiesAliveandentitiesUnreadable, and a verdict resting on "none are left" requires a complete census.Rescaled bounds
Scale bounds moved because the real breaking points did. All values are 1.6.0 → 1.7.0:
- TNT: base 100 → 40 per zone, ceiling 10 000 → 5 000 (hard cap 50 000 → 25 000).
- Chunks: 512 → 2 000, and the unit changes from chunks per tick to chunks per second — measured, the real limit sits around 160–180 chunks/s, so a per-tick ceiling was unrelated to the test's physics.
- Entities: 20 000 → 80 000.
REFERENCE_VALUES.stressTntis re-anchored on the results service in the same move: scoring a 1.6.0 report with the 1.7.0 reference, or the reverse, gives an incoherent score.Quieter fixes
blocksDestroyedcounted a ring of blocks nothing had destroyed: the before-count used a hardcoded radius of 30 and the after-count the platform's real half-extent, so at 150 TNT it compared 61×61 columns against 51×51 — and the error grew with the load.explosionImpactwas "20.0 minus minimum TPS" with a minimum that falls back to 20.0 when nothing was measured, so it published0.00, the best possible result, on an absence of measurement.- Protection plugins are no longer accused on an absence of evidence.
blockDamageSuppressedwas raised on "no survivors and no damage", which a never-ticked charge satisfies exactly as well as a neutralised explosion. It now requires a detonation actually observed, and the report names the subscribers toEntityExplodeEvent,BlockExplodeEventandExplosionPrimeEventas suspects, never as a verdict. - Forced chunks could outlive the run, the plugin and a restart, since force-load
is persisted into world data. The release path is fixed and a failed un-force
is now logged instead of swallowed — watch for
could not un-force chunk X,Y — it may stay loaded after restartin console. - The explosion probe survived
/bench stopand the server's lifetime, keeping three handlers live on every real player's explosion and holding the dead test instance with its thousands of charges. The standard/bench startpath also discarded every probe counter it collected, andtntCount/tntTotalwere read by the site's scoring (a difficulty factor up to ×1.05) but published by nobody, so the bonus could never apply. workCompleted: 0on a test that had done everything: work was derived from the entities alive at the end, and a test with transient entities has none. On Folia,blockPhysicspublished 16 000 tracked entities, 349 blocks/s and 20.0 TPS alongsideworkCompleted: 0. Work is now anchored on the entities the test actually created.- An entity that finished its life is no longer counted as frozen — a falling
block lands and becomes a block again.
blockPhysicsdeclared 20 % of its batch frozen (0.80 Paper, 0.77 Folia); after the fix, Paper readsentityTickRatio1.00 with no flag andworkCompleted16 000. - Tier windows are now constant. A per-tick quota floored at 1 made the window
10 s at 100 TNT but 5.5 s at 10, leaving four usable one-second slices, one of
them the salvo —
sampleCount23 against 37–40 everywhere else. - Hopper and redstone scales: early skip fired during preloading (Paper climbed to its 40 000-line ceiling with MSPT never above 3.9 ms) and the absorption criterion broke the scale at tier zero on an idle server.
/bench testrefused 8 of the 26 identifiers it advertises, and/bench stopdid not stop the run and left the server modified afterwards.- Chunk preloading froze the server for fifteen seconds on a first run at one host; the chunk scale now counts chunks delivered rather than requests sent.
- The local fallback report — the one that remains when submission fails — was empty or lost, and carried none of its statistics. The discarded-sample share is now published along with its reason, and two messages that sent operators hunting a submission failure that did not exist are gone.
- The stress scale and the report interface no longer answer in French regardless of your locale: unit and type names are protocol values, translated only for display.
- Non-gameplay tests were judged on the global region, idle by design; a region freeze is now counted in real seconds rather than as a single degraded second.
- The Folia exemption from the ticking-chunk wait made Folia score better than
Paper on identical hardware while that wait was still inside
durationSeconds. The wait now sits outside the window on every platform.
For extension authors
An extension's declared permission was dropped and
getTestPermissionanswered "no permission required" for every non-built-in identifier, so extension tests ran without their permission and without their parameter bounds. After updating, an extension test a non-op could previously launch may be refused — and one meant to be gated is now genuinely gated.Nine
protectedmembers leftAbstractBenchmarkTest's surface —checkMemoryBudget,getTotalDurationTicks,getWarmupSeconds,hasReachedDuration,isPositionLocked,loadChunksSquare,loadChunksSquareSpread,performAggressiveGc,unlockPlayerPosition— some deleted as callerless, the rest restricted toprivate; eight new ones replace them (armTickWitness,captureInProgressMetrics,freezeRegionMeasurement,publishWork,putIfMeasured,registeredChunkCount,tickingLease,witnessEntities). An extension built against 1.6.0 needs a recompile. That class is not part offr.wasabii.voxelBench.api.**, which is unchanged — hence the minor bump.Finally,
flags.scoreis no longer sent in the stress-limit payload. The plugin's own computed score left the wire; scoring belongs to the backend. Any third-party consumer reading that field loses it. No other payload field was removed.Upgrading
Drop the new jar in
plugins/and restart. There is no database migration, and your existingconfig.ymlkeeps working. Three points deserve attention:- The two new config keys will not appear in an existing
config.yml.config-versionis unchanged at 8 and the plugin does not migrate configs, so a server upgrading in place keeps its 1.6.0 file. Both keys default to their new behaviour in code (raise-host-tnt-quota: true,stress-limit.warm-start.enabled: true), so the new behaviour is what you get. It is to turn either one off that you must add the block by hand, or regenerate the file. - Delete
data/stress-warmstart.ymlbefore your first campaign, or wait outmax-age-days(30 by default). Remembered values at or above the current ceiling are rejected outright, but a warm start onto a pre-1.7.0 value wastes tiers. - Re-run your reference campaigns, for the reasons given at the top.
Compatibility
Java 16 or later, for Minecraft 1.17 through the latest release including 26.x. Tested on Spigot, Paper and Folia.
- 1.7.0stablearchivedAug 12, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1, 26.2 · 1.6 MB · 15 ↓No longer available
This release fixes what the bench was measuring, not how it scores. Three defects compounded, and together they meant the TNT stress scale never broke on any machine: it climbed to its configured ceiling on every host, measuring entity spawning and platform construction rather than your server. The headline is short — the bench now proves that a test ran before it publishes that test's number.
The explosion test detonated nothing: forcing a chunk loads it but never brings it to
ENTITY_TICKING, so fuses sat at 40/40 after ten seconds while the charges were reported "alive". Underneath, the measurement layer failed open — sixteen fallbacks across the region monitor and the tier verdict all meant "healthy server", so a tier that measured nothing was published as20.00 TPS / 0.00 ms, the best possible result. And under that,max-tnt-per-tick: 100, the default on Spigot and Paper, silently capped how much load ever reached the server.protocolVersionstays 1.2.0 andAPI_VERSIONstays 1.2.0: the wire contract is additive. What changed is what the numbers mean.Your previous results do not carry over
Benchmark and stress-limit results from 1.6.0 and earlier are not shifted. For
stressTnt(every platform) andstressHoppers(Paper, Spigot) they are void, and no Folia stress-limit result existed at all before this version. There is nothing to convert and nothing to re-read: no 1.6.0 report carries a single validity witness, so it is impossible to tell after the fact which of these causes applied to it.Re-run your campaigns, and expect lower limits. That is the measurement becoming real, not your machine regressing. Scale bounds moved as well, so a tier number and a load value no longer mean what they meant.
Measured breaking points on the corrected TNT scale, for reference: 1938 ±76 (Folia), 2229 ±92 (Paper), 2279 ±65 (Spigot). The corrected hopper scale breaks at 3112 ±71 on Paper, where it previously returned 40 000 — against 121 on Spigot and 2 661 on Folia, three mutually incompatible answers for one test.
Reports produced by 1.7.0 pre-releases are not comparable to this build either. Reject any that show
entityTickRatioof exactly1.00on a test with transient entities, orworkCompletedof exactly0; both were wrong by construction until late in the cycle.The explosion test now actually explodes
Forcing a chunk attests to loading, not to ticking, and only a ticking entity burns a fuse. Charges were created, accepted by the server, then frozen — and since frozen work costs a region nothing, TPS held at 20.00 while the scale climbed. Zones now hold their own ticking lease, released zone by zone (one zone finishing used to release the force-load of the other three, which stopped ticking mid-flight), and the measurement window does not open until the zones really tick. At tier 0, 498 blocks now fall where zero fell before.
On Paper that promotion takes about 20 seconds — measured, 9 → 200/200 chunks in 17.8 s on a 150×8 run, against 1.8 ms on Spigot. The wait is announced in console every 5 s, on the scoreboard and in the action bar, capped at 30 s, and counted outside the measurement window: it does not enter
durationSeconds,blocksPerSecondor the TPS/MSPT window. It is not a freeze.The gap between "loaded" and "ticking entities" is now measured rather than assumed. Over 8 zones and 200 forced chunks, Paper returned 88
ENTITY_TICKINGagainst 112TICKING— a level that runs blocks and random ticks, so it has every appearance of an active zone — while Spigot returned 200/200.A tier that measured nothing is no longer the best tier
The "no result" branch emitted 20.00 TPS and 0.00 ms across fourteen fields with
measurementDegraded: false: the most degraded case of all passed for the only healthy one. Symmetrically,loadAbsorbedRatio: -1— meaning "no witness" — was remapped to 1.0, so the tier was not saturated and it passed, word for word what that method's own documentation said it existed to prevent.The verdict now distinguishes "no witness" from "a witness saying we cannot judge" and returns
NOT_MEASURED— neither held nor broken — read before the health branch. A tier whose load was never censused therefore stops being labelled "health score low" in a message stored in the database and shown to you, as if your machine were at fault.Two consequences follow:
- A tier the server could not detonate within its window is now a broken tier. Explosion saturation was invisible to every other metric: queued work costs a region nothing, so TPS held at 20.00 and the degraded-second fraction at 0.00 across thirty-two tiers while four charges in five were never touched (798 detonated + 3202 still alive = 4000 requested). The tier now breaks below 95 % presented load, and the break reason names it.
- A collapsed region publishes its own collapse. The witness sampler is
scheduled in ticks of the measured region, so a frozen region never runs it and
the measurement vanishes exactly where it matters. A scheduled task that never
started in D seconds proves the region executed fewer than 20 ticks over that
span, which bounds its rate from above with no sampler at all. Published as
regionTaskStarvedandregionTickCeiling— on mobAI, under 0.68 tps, corroborating the 0.30 tps the region monitor measured independently.
Your server's own TNT cap no longer caps the run in silence
max-tnt-per-tick: 100is the default on Spigot and Paper. A test posting 1200 charges never saw most of them tick, so it measured a configuration file rather than your machine — silently, on almost every host.The cap is now raised in memory for the duration of the test and restored on every exit path, including
/bench stop, a timeout and an exception. Yourspigot.ymlis never written, and a restart re-reads it intact. Measured on 150×8 over two consecutive runs per platform, the second run proving the first restored:tntNeverTicked300 → 0,tntPresented900 → 1200/1200,detonationRatio0.89 → 1.00,loadAbsorbedRatio-1 → 1.00.Reports carry
hostMaxTntPerTick,hostTntQuotaSourceand one ofhostTntQuotaRaised/hostTntQuotaLimited/hostTntQuotaUnknown, so two runs of the same server stay distinguishable. Setbenchmark-tests.explosion.raise-host-tnt-quota: falseto refuse any in-memory change; the run is then reported as partial and a console warning names the value to raise yourself.The read itself was broken too.
Bukkit.spigot().getClass().getMethod(...)resolves an override declared by a non-public anonymous class, soinvokethrewIllegalAccessExceptionon every call, swallowed by a catch — and every read fell back tonew File("spigot.yml"), relative to the JVM's working directory, missing as soon as a panel or a systemd unit sets another one. There are three ordered routes now, and the config-shape check refuses any file that is not aspigot.yml, because reading a falsely high cap elsewhere would credit a host with absorbing load it never received.Every report now carries the evidence that it measured something
Each test publishes a witness saying whether its entities actually moved, how many of its chunks really ticked entities, and how much of the announced load the server was actually given. New fields include
entityTickRatio,tntPresented,tntNeverTicked,detonationRatio,detonationUnaccounted,explodeForeign,zoneChunksEntityTicking,workCompletedwith its unit,loadAbsorbedRatio,sampleCountand per-tiermeasurementDegraded.systemInfonow also distinguishes the host machine from what the server can actually use, throughavailableProcessors,memoryLimitGbandprimaryDisk.usableGb/.totalFsGb.entityTickRatiois built onticksLived, the one counter that only advances when an entity is ticked. It is armed on every test, sampled over at most 256 entities and collected once a second from each owning region, so a partial freeze stays visible;entityTickSamplesandentityTickConclusivesit beside it, because "the sampler never ran" and "nothing was readable" are different failures. For explosions the fuse is the exact witness — it decrements once per entity tick, so a charge still at its initial value was never ticked once. The estimate it replaces returned 694 charges presented against 800 detonated, a ratio of 1.15 that novalue < thresholdguard could ever catch.Two rules follow for anyone reading a report:
- Read the flags before the numbers.
loadNotPresented,loadPresentationUnmeasured,zoneTickingNotReady,zoneTickingCapped,entitiesNeverTicked,entitiesPartiallyTicked,entityCensusPartial,regionTaskStarvedandmeasurementDegradeddescribe a doubtful measurement, not a slow machine, and that tier does not compare to the others. They are raised only when something is wrong: a healthy report carries none, because a permanent flag is noise and noise stops being read. - Never read
-1as0. It means "cannot conclude", never "nothing moved".
On a populated server, TNT lit by your players or your redstone is now counted separately as
explodeForeigninstead of being credited to the machine under test. The old probe incremented for anyTNTPrimedexplosion anywhere on the server, and that counter decides the tier verdict: on the same host the same day, 1 zone gave 1.00, 8 zones gave 0.67, and 8 zones with simulation distance raised gave 0.48 with no survivors — three limits for one machine, driven by a bench parameter. Still, run your benchmarks on an idle server: foreign detonations no longer distort the verdict, but they contaminate the run.Warm start, and a break that has to repeat
The last stable value of each stress test is remembered in
data/stress-warmstart.yml, and the next run starts near 70 % of it instead of climbing the whole scale again. If the warm probe breaks, the runner refines downwards; if nothing there holds, it falls back to a clean full climb. Stale and other-algorithm entries are ignored. Tunable understress-limit.warm-start:enabled(true),factor(0.70),max-refine-down-steps(6),max-age-days(30).Warm start makes a run depend on the previous run of that machine, so two successive runs of the same server are no longer directly comparable with it on. Set
stress-limit.warm-start.enabled: falsewhen you are comparing two servers against each other.Separately, a broken tier is now replayed before the scale stops, on the way up and during refinement, with the confirmation variance tightened from 30 % to 12 %. An isolated hiccup no longer condemns a healthy machine. Runs take longer, and some limits go up relative to 1.6.0 for this reason.
/bench tier — walk one scale on its own
/bench tier <test> [tiers=N] [from=N] [to=N] [duration=SEC] [zones=N] [cold] [noskip]runs a single stress test up its scale to diagnose it, without launching a full campaign. It sits behind the newvoxelbench.tierpermission (op by default), never writes the warm-start memory, and submits on the unit-test channel governed byreports.backend.unit-tests— never as a stress-limit report. Note thatfromandtoare loads, not tier numbers.Folia
Headline TPS and MSPT now describe the measured window instead of being an exponential-moving-average snapshot read at the worst moment of the run, so every Folia headline number changes magnitude.
An entity the current thread cannot read is also no longer counted as a dead one. Reading an entity from a thread that does not own its region throws under Folia, and the old code turned every missed read into a corpse — which produced "no charges left", then "so they exploded", then "so a plugin is cancelling the damage": three conclusions stacked on missed reads.
EntityCensusnow publishesentitiesTracked,entitiesAliveandentitiesUnreadable, and a verdict resting on "none are left" requires a complete census.Rescaled bounds
Scale bounds moved because the real breaking points did. All values are 1.6.0 → 1.7.0:
- TNT: base 100 → 40 per zone, ceiling 10 000 → 5 000 (hard cap 50 000 → 25 000).
- Chunks: 512 → 2 000, and the unit changes from chunks per tick to chunks per second — measured, the real limit sits around 160–180 chunks/s, so a per-tick ceiling was unrelated to the test's physics.
- Entities: 20 000 → 80 000.
REFERENCE_VALUES.stressTntis re-anchored on the results service in the same move: scoring a 1.6.0 report with the 1.7.0 reference, or the reverse, gives an incoherent score.Quieter fixes
blocksDestroyedcounted a ring of blocks nothing had destroyed: the before-count used a hardcoded radius of 30 and the after-count the platform's real half-extent, so at 150 TNT it compared 61×61 columns against 51×51 — and the error grew with the load.explosionImpactwas "20.0 minus minimum TPS" with a minimum that falls back to 20.0 when nothing was measured, so it published0.00, the best possible result, on an absence of measurement.- Protection plugins are no longer accused on an absence of evidence.
blockDamageSuppressedwas raised on "no survivors and no damage", which a never-ticked charge satisfies exactly as well as a neutralised explosion. It now requires a detonation actually observed, and the report names the subscribers toEntityExplodeEvent,BlockExplodeEventandExplosionPrimeEventas suspects, never as a verdict. - Forced chunks could outlive the run, the plugin and a restart, since force-load
is persisted into world data. The release path is fixed and a failed un-force
is now logged instead of swallowed — watch for
could not un-force chunk X,Y — it may stay loaded after restartin console. - The explosion probe survived
/bench stopand the server's lifetime, keeping three handlers live on every real player's explosion and holding the dead test instance with its thousands of charges. The standard/bench startpath also discarded every probe counter it collected, andtntCount/tntTotalwere read by the site's scoring (a difficulty factor up to ×1.05) but published by nobody, so the bonus could never apply. workCompleted: 0on a test that had done everything: work was derived from the entities alive at the end, and a test with transient entities has none. On Folia,blockPhysicspublished 16 000 tracked entities, 349 blocks/s and 20.0 TPS alongsideworkCompleted: 0. Work is now anchored on the entities the test actually created.- An entity that finished its life is no longer counted as frozen — a falling
block lands and becomes a block again.
blockPhysicsdeclared 20 % of its batch frozen (0.80 Paper, 0.77 Folia); after the fix, Paper readsentityTickRatio1.00 with no flag andworkCompleted16 000. - Tier windows are now constant. A per-tick quota floored at 1 made the window
10 s at 100 TNT but 5.5 s at 10, leaving four usable one-second slices, one of
them the salvo —
sampleCount23 against 37–40 everywhere else. - Hopper and redstone scales: early skip fired during preloading (Paper climbed to its 40 000-line ceiling with MSPT never above 3.9 ms) and the absorption criterion broke the scale at tier zero on an idle server.
/bench testrefused 8 of the 26 identifiers it advertises, and/bench stopdid not stop the run and left the server modified afterwards.- Chunk preloading froze the server for fifteen seconds on a first run at one host; the chunk scale now counts chunks delivered rather than requests sent.
- The local fallback report — the one that remains when submission fails — was empty or lost, and carried none of its statistics. The discarded-sample share is now published along with its reason, and two messages that sent operators hunting a submission failure that did not exist are gone.
- The stress scale and the report interface no longer answer in French regardless of your locale: unit and type names are protocol values, translated only for display.
- Non-gameplay tests were judged on the global region, idle by design; a region freeze is now counted in real seconds rather than as a single degraded second.
- The Folia exemption from the ticking-chunk wait made Folia score better than
Paper on identical hardware while that wait was still inside
durationSeconds. The wait now sits outside the window on every platform.
For extension authors
An extension's declared permission was dropped and
getTestPermissionanswered "no permission required" for every non-built-in identifier, so extension tests ran without their permission and without their parameter bounds. After updating, an extension test a non-op could previously launch may be refused — and one meant to be gated is now genuinely gated.Nine
protectedmembers leftAbstractBenchmarkTest's surface —checkMemoryBudget,getTotalDurationTicks,getWarmupSeconds,hasReachedDuration,isPositionLocked,loadChunksSquare,loadChunksSquareSpread,performAggressiveGc,unlockPlayerPosition— some deleted as callerless, the rest restricted toprivate; eight new ones replace them (armTickWitness,captureInProgressMetrics,freezeRegionMeasurement,publishWork,putIfMeasured,registeredChunkCount,tickingLease,witnessEntities). An extension built against 1.6.0 needs a recompile. That class is not part offr.wasabii.voxelBench.api.**, which is unchanged — hence the minor bump.Finally,
flags.scoreis no longer sent in the stress-limit payload. The plugin's own computed score left the wire; scoring belongs to the backend. Any third-party consumer reading that field loses it. No other payload field was removed.Upgrading
Drop the new jar in
plugins/and restart. There is no database migration, and your existingconfig.ymlkeeps working. Three points deserve attention:- The two new config keys will not appear in an existing
config.yml.config-versionis unchanged at 8 and the plugin does not migrate configs, so a server upgrading in place keeps its 1.6.0 file. Both keys default to their new behaviour in code (raise-host-tnt-quota: true,stress-limit.warm-start.enabled: true), so the new behaviour is what you get. It is to turn either one off that you must add the block by hand, or regenerate the file. - Delete
data/stress-warmstart.ymlbefore your first campaign, or wait outmax-age-days(30 by default). Remembered values at or above the current ceiling are rejected outright, but a warm start onto a pre-1.7.0 value wastes tiers. - Re-run your reference campaigns, for the reasons given at the top.
Compatibility
Java 16 or later, for Minecraft 1.17 through the latest release including 26.x. Tested on Spigot, Paper and Folia.
- 1.6.0stablearchivedJul 6, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1, 26.2 · 1.5 MB · 29 ↓No longer available
This is a hardening and tooling release built on top of the Folia work shipped in 1.5.0. The headline is that Folia measurement is now accurate and stable across the whole suite: a silent task-cancellation bug that produced ghost timeouts, a stuck action bar and garbage TPS deviation is fixed, and every gameplay test — built-in and extension alike — is now measured on the region actually under load. The release also brings a rebuilt monitor subsystem, a new connectivity diagnostic command, and a large internal cleanup.
Spigot and Paper runs remain byte-identical to 1.5.0. Every change described below is gated behind
ServerCompat, so if you do not run Folia, your numbers are exactly the ones 1.5.0 produced.Folia timers that were never actually cancelled
On Folia,
ScheduledTask.cancel()threw anIllegalAccessExceptionbecause the task class is not public, and the exception was silently swallowed. The cancellation looked like it had happened; nothing had. Timers, test timeouts and the per-region samplers all kept running.The symptoms did not look related to one another:
- Phantom "test timeout" messages arriving in chat after a run had already finished.
- A "Multi-Core" action bar that stayed on screen with nothing behind it.
- Nonsensical TPS and MSPT standard deviation, produced by stale samplers piling up on top of one another.
The fix is
setAccessibleon the task class, so cancellation now takes effect. Alongside it, all internal timers are tracked so they can be cancelled as a group, and Folia chunk loading is now performed per region.Every gameplay test is measured on the region under load
On Folia, the global region is idle while a benchmark runs — the work is happening in whichever region the test occupies. Until now, only the six dispersed tests fed the shared per-region monitor. Every other test read its TPS and MSPT from the idle global region, which carries none of the load.
The non-dispersed gameplay tests now feed that same monitor: mob AI, block physics, bone-meal, redstone, liquids, combat, villagers, collisions, cramming, projectiles, pathfinding and chunk ticking. Extension tests are covered too, and get correct Folia measurement without any change on their side.
/bench pingtells you where the connection breaksWhen a report will not submit, the useful question is which layer is failing.
/bench pingruns a staged probe — DNS, then TCP, then TLS, then HTTP — and reports pass or fail for each step: an unresolved host, a blocked port, a TLS handshake failure, a timeout, or the HTTP status when the service answers.The probe runs off the main thread and is available on every build.
Monitor subsystem rewrite
The monitor subsystem was restructured with no change in behaviour. The boss-bar logic moved into its own manager, the web dashboard's HTML was externalized with Chart.js bundled locally, and the monolithic monitor command was split into per-subsystem handlers. The two largest files shrank by roughly two to three times.
Quieter fixes
- The spurious "Scoreboard created but NOT active" warning is gone. It was racing the asynchronous scoreboard assignment and reporting a failure that had not occurred.
/bench verifyno longer crashes the scheduler inSlpInjector, and no longer floods chat on a successful link.- The temporary benchmark world is written with explicit flat layers, and the
loadChunksSpreadcompletion latch is hardened against a missed signal on Folia. - Run-to-run variance is lower: allocations were hoisted out of the measured sampling window across the gameplay tests.
- Internal cleanup from a full code audit — the standard and custom autobot runners are unified, ranged arguments go through a single shared parser, 309 unused translation keys were removed, and assorted dead code and leaks were cleared out.
Compatibility and upgrade
Java 16 or newer, Minecraft 1.17 through the latest release including 26.x. Tested on Spigot, Paper and Folia. The continuous-integration compatibility matrix now extends to Minecraft 26.2 across Paper, Spigot and hybrid servers, alongside the versions already covered.
Drop the new jar into
plugins/and restart. There is no configuration change and no database migration. - 1.5.0stablearchivedJul 5, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1 · 1.4 MB · 0 ↓No longer available
This release makes Folia a platform VoxelBench can actually measure.
/bench startnow runs the entire test suite on Folia — no skipped tests, no failing ones — and produces a complete report instead of a partial one. The numbers in that report are measured where Folia does its work, and the report says so, so nothing gets compared apples-to-oranges against a single-threaded run.Spigot and Paper runs are unchanged — byte-identical scores to 1.4.0. Every Folia behaviour is gated behind server-type detection. If you do not run Folia, upgrading changes nothing about your results.
The whole catalogue runs on Folia
Until now, a Folia benchmark came back with holes in it. The entire built-in catalogue now executes correctly on Folia's per-region scheduler: chunk loading, mob spawning, redstone, block physics, chunk ticking, hoppers, explosions, lighting, tile entities, mob AI, bone-meal growth, liquids and the rest.
The reason it works is that the benchmark is region-aware end to end, not merely patched test by test. World acquisition, weather and time, the player's teleport, gamemode and state, chunk loading, force-loading and cleanup all run on the correct region thread, so a run no longer commits main-thread or async-region violations and no longer trips the server watchdog. Force-loading in particular now goes through plugin chunk tickets rather than
setForceLoaded, which only ever addressed the global region. Shared helpers —ZoneExecutor, a region-awareBenchmarkContext,AbstractBenchmarkTest— carry that logic so it is not re-implemented, differently, in every test.TPS and MSPT measured where the work happens
On Folia the work is spread across region threads while the global region sits idle. A single global TPS/MSPT reading therefore measures the one thread that is doing nothing — a number that looks like a result and is meaningless.
The benchmark now samples TPS and MSPT per region and aggregates them into the same result fields you already know — the average across regions, with minimum and maximum standing for the best and worst region. Reports also carry a
regionizedflag and a region count in their envelope, so the service reading them knows it is looking at a multi-threaded run and interprets the figures accordingly.Alongside this, select tests — block physics and mob spawning — now also report throughput, expressed as work per second, which suits a parallel model better than a per-tick figure. These metrics are added, never substituted: the existing ones are untouched, and a service that does not know the new ones falls back to what it already understood.
Fixes
- Gamemode race on state restore. Restoring the player could throw
Cannot set gamemode asyncon Folia. The player is now teleported first, and the gamemode restored only once the asynchronous teleport has settled. - Temporary worlds could not be cleaned up. Unloading a world at runtime is
unsupported on Folia and threw
UnsupportedOperationException. The temporary benchmark world is now left in place and removed cleanly on the next server start, instead of erroring at the end of every run. - Stale timeout lines. Tests that had already finished could still log a "test timeout" afterwards. The timeout is now cancelled synchronously the moment the test completes, so it cannot fire late on a backed-up global region.
- Leaked internal timers.
ScheduledTask.isCancelled()always answered "cancelled" on Folia — it reflected a method Folia tasks do not expose — so everyif (!isCancelled()) cancel()quietly skipped the cancel. Samplers and timeouts piled up across a run, raising per-tick overhead and GC pressure, and degrading the very measurement they served. They are now cancelled correctly. - Chunk-ticking build load is throttled on Folia so it no longer floods a single region.
Extension API
The extension API moves to 1.2.0, with a region-aware
BenchmarkContext. The change is backward compatible, and the API remains PREVIEW.It lets a third-party plugin register its own benchmark tests against VoxelBench, with typed parameters, declared output metrics, full lifecycle hooks and mock helpers for unit tests. It is functionally complete and exercised by the bundled sample plugins, but it is not officially announced for public use yet — expect changes until the public release.
Upgrading
Drop the new jar in
plugins/and restart. There is no configuration change and no database migration.One thing to expect on Folia: a faithful Folia benchmark deliberately disperses its work across region threads. It shines on machines with many cores, and is intentionally slower on a server with only a few.
Compatibility
Java 16 or later for MC 1.17 through 1.21.x; Java 21 recommended for MC 1.20.5 and above; Java 25 LTS for MC 26.1.x. Tested on Paper, Purpur, Pufferfish, Folia and Spigot. Continuous integration now includes Folia boot and runtime-compat smoke tests on Folia 26.1.2, with assertions that catch thread and region violations.
- Gamemode race on state restore. Restoring the player could throw
- 1.4.0stablearchivedJun 17, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1 · 1.4 MB · 18 ↓No longer available
This is a tooling and interface release. The stress-limit mode was rebuilt so a run keeps climbing until the server genuinely buckles, the CPU tests let you choose which part of the processor you stress, the
/benchmenu now builds itself from the live test catalogue, tests can be configured in-game through a real form, and reports carry an inventory of the server's environment so the leaderboards can flag setups that inflate scores.No breaking changes to standard-benchmark scoring.
Stress-limit, rebuilt
The stress-limit mode used to plateau early, or stop on a reading that was never a real rupture. It now ramps its tiers adaptively and measures lag faithfully, so a run continues until something actually gives.
- Faithful TPS/MSPT sampling, and the standard load clamp is bypassed so each tier genuinely scales up instead of being held back by a limit meant for the standard benchmark.
- Per-type starting loads calibrated from in-game testing.
- Bounded measurement window for the hopper vector, and an inter-tier safety sweep of shared zones between two tiers.
- A shared report envelope for stress-limit results.
Three failure modes are gone with it: false ruptures caused by a mean-based break trigger, endless hopper tiers, and tiers that never reached the breaking point at all.
A seventh vector: villager brain
The new villager-brain vector piles on villager AI and pathfinding until the brain ticks drown the server. That brings the count to seven stress vectors.
Custom stress profiles
You can now compose your own stress recipe — which vector, which starting load, which ramp — in a YAML file dropped into
plugins/VoxelBench/custom_benchmarks/, then run it from chat or from the menu. The GUI carries a custom-profile browser so those files can be launched without typing their names.CPU compute kernels
The single-core and multi-core tests can run one of four compute kernels, so you can target the part of the CPU you care about instead of a single fixed workload:
int— integer ALU throughputfloat— floating-point throughputmemory— memory-bound access patternbranch— branch-prediction stress
The kernel is selectable per test, through the command parameter, tab-completion, or the in-game configurator.
The menu is driven by the test catalogue
The
/benchGUI was largely rewritten. Test menus now build themselves from the live test catalogue, which means the menu can no longer drift out of sync with what the plugin is actually able to run: new built-in tests, and third-party tests registered through the extension API, appear on their own with no menu edits.The overhaul also brings a Mods screen, a server config-profile screen, a dynamic stress menu, a clearer tab bar, and the removal of a duplicate "Tests" entry in the main menu. All the new screens, and the test parameter labels, are localized in English and French.
One defect went with it:
refresh()blanked the screen, which broke pagination, value cycling and every refreshed menu. It now rebuilds and re-opens correctly.Configuring a test in-game
On Paper 1.21.6+, opening a test brings up a proper configuration window instead of asking you to remember its parameters. Each test gets one form, with dropdowns for choice parameters, readable and translated labels, and free-text entry via shift-click.
This runs on a new reusable Paper dialog layer (
gui.dialog). On any other platform the configurator falls back automatically to an anvil, chat or inventory flow, so every server can still set parameters cleanly.Tab-completion suggests values
Built-in tests now suggest sensible values for each parameter as you type, not just the parameter names. Tests added through the extension API get the same value completion via
ParamSpecsuggested values.Environment flagging
Benchmark reports now inventory the server's mods, plugins and configuration profile. This is what lets the leaderboards spot and flag setups that inflate scores — tweaked view-distance, spawn limits, performance mods — so honest runs are not drowned out by doctored ones. The plugin only reports the data; the flagging happens server-side.
Smaller fixes
- The explosion test's proportional TNT spawn is now gated to stress mode only, so the standard benchmark is unaffected by it.
- The pre-flight memory budget anticipates a test's peak allocation rather than its starting one.
Extension API
The extension API moves to 1.1.0 — suggested values, backward compatible, and still marked PREVIEW.
Third-party plugins can register their own benchmark tests against VoxelBench, with typed parameters, declared output metrics, full lifecycle hooks and mock helpers for unit tests. The surface is functionally complete and exercised by the bundled sample plugins, but it is not officially announced for public use yet: expect changes until the public release.
Upgrading
Drop the new jar in
plugins/and restart. There is no config or database migration — the new options ship with sensible defaults. On Paper 1.21.6+ the in-game configuration windows light up on their own; every other platform keeps the anvil / chat / inventory flow. Custom stress profiles live next to the benchmark profiles inplugins/VoxelBench/custom_benchmarks/; delete the samples you do not need.Java and compatibility
Java 16+ for MC 1.17 to 1.21.x. Java 21 recommended for MC 1.20.5+. Java 25 LTS for MC 26.1.x.
Tested on Paper, Purpur, Pufferfish, Folia and Spigot. The in-game configuration windows require Paper 1.21.6+; everything else degrades gracefully on older or non-Paper platforms.
- 1.3.0stablearchivedMay 31, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1 · 1.3 MB · 14 ↓No longer available
This release is about not bringing your server down while measuring it, knowing what kind of machine a measurement came from, and being told what could spoil a run before it starts. Three layered guards now stand between the bench and an
OutOfMemoryError, the plugin identifies the host it runs on, and the confirmation before a run has become a screen listing everything that could spoil the result.Chunk-loading results are not comparable with 1.2.1 and earlier. The standard test now loads a different number of chunks; the details are in the next section.
Chunk loading: 16 000 chunks instead of 40 000
The standard
chunkLoadingtest loads 2000 chunks per zone instead of 5000. With the default eight dispersed zones, a run covers 16 000 chunks instead of roughly 40 000 — a figure that was itself clamped to about 37 800 by the adaptive memory cap on a 6 GB heap.The reason is the memory work below. Forty thousand chunks routinely pushed heap to the 92 % mark that now trips the watchdog on 6 GB hosts; sixteen thousand stays comfortably under the threshold at which the pre-flight budget refuses to start a test at all. The bundled
standard.ymlprofile mirror was updated to match, andlow-memory.ymlwas taken down to 1000 chunks per zone across four zones — 4000 total — so that it stays meaningfully lighter than the new standard.The consequence for you: a 1.3.0 run is not directly comparable to a 1.2.1 run for this test, nor for the aggregate score that weights it. Leaderboards for chunk loading have to be read per plugin version.
Three guards against running out of memory
A bench that would previously have taken the server down now refuses to start, or stops itself, and says which of the two happened.
- After each test — cleanup triggers an aggressive collection (three cycles
with a stabilisation check), recovering the 10 to 15 % of peak heap that each
test leaves behind as unreclaimed garbage. Toggle with
benchmark.post-cleanup-gc, on by default. - Before each test — a memory budget reads heap usage. Below 75 % the test
proceeds; between 75 and 85 % it logs a warning and forces a collection; at
85 % or above it refuses to start and the test is returned as
SKIPPEDwith the reasonmemoryPressure. Tunable withbenchmark.memory-budget.{warn-pct,fail-pct}. - During each test — an asynchronous heap poller. When heap stays at or
above 92 % for three consecutive seconds, it fires a clean abort: the test is
reported as
FAILUREwith anoom_abortwarning and cleanup is forced, instead of letting the JVM throw an actualOutOfMemoryError. Tunable withbenchmark.memory-watchdog.{critical-pct,hold-seconds}.
The report records what kind of host it came from
Five probes look for signs of a free-tier host: marker files (Aternos, Minehut, FalixNodes, Server.pro, the MSH wrapper, Pterodactyl), plugin names, environment variables, JVM resources (a heap of 1 or 2 GB, a single-CPU JVM), and
/proc/1/cgroupmarkers. Their verdicts aggregate into a five-level tier —DEDICATED_OR_VPS,LIKELY_PAID_HOSTED,SUSPECTED_FREE,LIKELY_FREE,CONFIRMED_FREE. A single vendor marker forces the tier to at leastLIKELY_FREEwhatever the score says.This surfaces twice. At
/bench start, a non-blocking chat warning names the detected provider, the signals that pointed at it, the variance you should expect, and points to/bench custom run free-host. And every report now carries ahostingEnvironmentblock — tier, score, vendor, signals, allocated heap, processor count — which is what the service needs to keep free-tier runs out of the public leaderboards. That block is deliberately left out of anonymization even inFULLmode: without the provider name the filtering cannot work.Detection can be turned off with
hosting-detection.enabled.Two profiles for hosts that cannot take the standard run
free-host.ymlis roughly a tenth of the standard workload, sized to fit under the free-tier killers on Aternos and MSH — TPS killer, RAM cap, CPU throttle.low-memory.ymlis roughly a third, for the 2 to 4 GB heaps of small VPS and dedicated hosts. Neither is comparable with/bench start;low-memory.ymlruns remain comparable with each other.The confirmation before a run is now a screen you can read
The old flow asked you to type
/bench starttwice within ten seconds, which told you nothing about why a second thought might be warranted. It is replaced by an inventory screen with one tile per detected risk, coloured by severity:- Critical — the target world is not a
voxelbench_*world; or automatic temporary worlds are disabled and no world is pinned. - Warning — the target world is not flat; other players are online and will
feel the lag; the server is already loaded before the bench (TPS below 19 or
MSPT above 30); the hosting tier is
LIKELY_FREEor worse. - Info — plugins likely to interfere, such as WorldGuard, GriefPrevention or Spark.
Clicking a tile prints the full details in chat. A critical finding disables the green Start button; only a Force button, visible to holders of
voxelbench.start.force, can launch through it. Closing the screen without clicking counts as a cancellation. When nothing at all is detected, the screen is skipped and the run starts.Disk and memory, measured rather than guessed
systemInfo.primaryDiskclassifies the device holding the world folder as NVMe, SSD, HDD, RAMDISK or unknown. The classification tries filesystem type first (tmpfsorramfsmeans a RAM disk), then the device name (nvme*), then the Linuxrotationalflag, and only then a heuristic on the model string. Model, type, size and mount point reach the report and appear as a dedicated tile in the disks screen.systemInfo.memoryexposes each memory module: capacity, speed, DDR type, manufacturer, part number and bank label, plus an aggregate type and speed across modules — reported as mixed when they differ. This comes from SMBIOS, which on Linux requires root, so a non-root server returns an empty list; it does so with stableUNKNOWNand-1sentinels rather than a different JSON shape. CAS latency and timings are not included and will not be: they need kernel-level SPD access that no JVM API offers.Anonymization was extended to cover both blocks. The mount point is always masked in partial and full modes, because
/home/<user>/...leaks a Linux username. The disk model is generalised to "Generic NVMe" and friends in full mode. Part numbers are hashed in partial mode and redacted in full — hashing rather than dropping keeps the "do these two servers have identical RAM kits?" comparison working. Manufacturer and bank label are redacted in full mode./bench info, rewritten from top to bottomFifteen sections, where several used to be near-empty titles. The existing
cpu,ram,disk,network,spigot,javaandsystemsections were enriched, and eight are new:performance(live TPS, MSPT and CPU),plugins(with enabled and disabled state),worlds(chunks, entities, players and seed per world),sensors(CPU temperature, voltage, fans),auth(mode, anonymization, server ID, rate limit),bench(last run summary),build(plugin version, profile, full JVM arguments) andhosting(tier and signals).Columns align to the pixel. A new width utility mixes bold and regular spaces to pad with 1-pixel precision, where plain spaces only offer a 4-pixel resolution. Tab completion now proposes the fifteen real sections instead of four obsolete ones. And
/bench info authmasks the server ID by default, printingabcd…wxyz (64 chars);/bench info auth hashreveals it.Chat says what actually happened
Until this release, every test ended on "✓ Test complete" regardless of whether the watchdog had killed it or the memory budget had refused to start it. The outcome line now branches on the real status: ✓ complete, ⊘ skipped with its reason, ✗ failed with its reason.
Wiring that up uncovered an older defect. The two-argument and eight-argument constructors of the unified test result set
successbut never synchronisedstatus, which stayed at its declared default of failure. Nobody read the status, so nothing showed — until the outcome line started reading it, at which point every successful test built through those constructors would have been displayed as failed. Both constructors now set the two fields in lockstep./bench stopnow stopsThe stop command was lying. Three coupled bugs, each hiding the next:
- Force-stopping did not run the active test's cleanup, so its
isStoppedflag was never flipped and the scoreboard stayed on screen. - Force-cleanup did not mark the test as finished, so when the asynchronous worker of a disk or network test eventually ground through its phases and reached the end, it fired its callback anyway — a stale "✓ Test complete" arriving seconds or minutes after you pressed stop.
- The manager did not check the stopped flag when that callback arrived, so even a fix to the second bug would have left the line printing.
All three are fixed. Cleanup is synchronous and the scoreboard disappears immediately, the late callback is suppressed, and no outcome line is emitted for a stopped run.
When the service refuses a report, you can see why
A 4xx response from the backend now surfaces its body in chat. Instead of
HTTP 400, you readHTTP 400 — missing required field 'X' at /tests[2]/status. Staging builds additionally write every response body next to the request payload inplugins/VoxelBench/raw-reports/, so the two can be diffed side by side; on production builds this does nothing.Upgrading
The first run after upgrading on a host with 4 GB of heap or less may show one or two skipped tests from the new memory budget. That is the intended behaviour: the bench is refusing to start tests that would have run the server out of memory. Check heap usage with
/bench info ram, or raise-Xmx.Reports now carry three new blocks —
hostingEnvironment,systemInfo.primaryDiskandsystemInfo.memory. A service that predates them ignores them silently; there is no schema break. - After each test — cleanup triggers an aggressive collection (three cycles
with a stabilisation check), recovering the 10 to 15 % of peak heap that each
test leaves behind as unreclaimed garbage. Toggle with
- 1.2.1betaarchivedMay 22, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1 · 1.3 MB · 4 ↓No longer available
VoxelBench 1.2.0 is, for the most part, a release for people who write benchmark tests rather than for people who run them. From a server operator's seat the plugin behaves as 1.1.0 did, with two exceptions that matter: hybrid runtimes finally generate and keep their benchmark worlds, and the
/benchcommand stops discarding what you type. The bulk of the diff is a new typed extension API for plugin developers who want to add their own tests — a surface that is in preview and not yet announced for public third-party use.Hybrid runtimes: worlds that generate flat, and survive a restart
On Arclight and Mohist, the flat benchmark world was not flat. Forge silently intercepts
WorldType.FLATon those runtimes, so the world VoxelBench asked for was not the world it received. A newFlatChunkGeneratorbypasses the NMS layer through Bukkit's ownChunkGeneratorAPI. It is activated only on hybrid runtimes: Paper, Spigot and Folia keep exactly the code path they had.On hybrids, the world also used to disappear on every server restart. Worlds an operator created (
voxelbench_<name>) are now loaded again when the plugin enables, so a benchmark world set up once stays set up.Related:
/bench world create <name>now refuses when a folder of that name is already on disk from a previous session. It used to load the legacy chunks and layer the new flat generation on top, which produced half-flat terrain.One test was affected badly enough to be worth naming. On hybrids,
world.spawnEntity()returns null without complaint when the target chunk is not ticking, and each of those nulls was counted as a successful spawn — sosamples.entitySpawn"finished" in zero seconds having spawned zero mobs. The test now pre-loads the nine chunks around the centre of its zone before the spawn loop starts.Continuous integration grew a hybrid smoketest matrix to keep these fixed: Mohist 1.20.1 (via a community mirror), Arclight 1.20.1 and Arclight 1.20.4, with Mohist 1.20.2 and NeoTenet 1.21.1 soft-skipped where upstream is broken.
The
/benchcommand says what it did, and keeps what you typedExtension tests now accept
key=valueparameters on the command line —/bench test myext.foo count=20 url=https://example.com. They were previously thrown away.The rest of the command surface was tightened in the same pass:
/bench test <id> <TAB>completes from the live registry, so extension tests appear where you would expect them./bench tests listis built from that registry too — grouped by category, each test tagged with its owner — instead of a hard-coded list./bench testsis now an alias of/bench test, in both the dispatcher and tab completion./bench helpis an explicit case rather than the silent default, and an unknown subcommand printsCMD_UNKNOWNbefore the help dump instead of looking like help was what you asked for./bench custom runrespects the world pinned with/bench world set, which it previously ignored.
Results are rendered as three distinct outcomes rather than two: success (green ✓), skipped (yellow ⊘) and failure (red ✗). The message attached to a failure is now read from the result's dedicated
errorfield; it was being looked up inmetrics["error"], which extension tests never wrote to, so failures explained themselves to nobody.Two smaller irritations went with it. Boolean metrics display "Yes" and "No" instead of
[Missing: result.no], and thecreateScoreboard() returned nullwarning was demoted to verbose — it is the normal, expected outcome for hardware tests and for every extension test, not an error worth a console line.Finally,
mobCountinsamples.entitySpawnwas silently clamped at 5000. The cap is now 50,000: asking for 10,000 mobs gets you 10,000 mobs.Extension tests no longer trigger phantom timeouts
An extension test that had completed normally could still trip the parent's safety timeout 60 or 120 seconds later, killing a run that had nothing wrong with it. The adapter now routes its result through
finishTest()instead of invoking the callback directly, which disarms the timeout on the way through.Extension API 1.1, in preview
The API 1.1 surface described below is marked
@ApiStatus(EXPERIMENTAL). It is usable and documented, but it is not yet announced for public third-party use, and the stability guarantees are the ones set out indocs/API_STABILITY.md§5.Tests declare their inputs and outputs, and the host checks them
ParamSpecgives a test a typed declaration of the parameters it accepts —intParam,longParam,doubleParam,stringParam,booleanParam, each with a default value, an optional range, a description and a required flag.MetricSpecis its symmetric counterpart for what the test emits:scalar,scalarInt,series,flagandtext, annotated with a unit, ahigherIsBetterdirection, a precision and aprimarymarker.Declaring is only half of it. The host now runs
ParamValidatorbefore callingrun(): a missing required parameter fails the test, a type mismatch is coerced or fails, and an out-of-range value is clamped with a warning to the sender. After a successful run, a warn-onlyMetricValidatorcompares what was emitted against what was declared, surfacing typos and drift without failing a run that otherwise worked.Because the host derives the legacy hint list from the declared specs,
TestDescriptor.Builder.paramHints(String...)is deprecated in favour of.params(ParamSpec...)."Not applicable" is now different from "failed"
TestResult.Statusis a tri-state enum —SUCCESS,FAILURE,SKIPPED— with aTestResult.skipped(reason)factory alongsidesuccess(...)andfailure(...). The distinction is between "the server lost" and "this test does not apply here", and it is meant to be honoured downstream: a backend should exclude skipped tests from leaderboards rather than count them as losses.Two footguns turned into compile errors and loud logs
TestBuilder.build(...)now receives aBuildContext, a narrow build-time view with no world, no sender and no stop signal;BenchmarkContextextends it for the runtime view, so runtime callers see no change. "ctx.getWorld()returns null at build time" is now a compile error rather than a null at the worst moment. Samples that never touched the context at build time are source-compatible; the ones that did were already buggy.TestCompletionwraps the rawConsumer<TestResult>callback with single-call enforcement (a second call logs SEVERE rather than doing something undefined), an automatic bounce to the main thread, and typedsuccess/failure/skippedshortcuts. The host carries its own anti-double-callback guard, whose SEVERE log names the offending extension.ApiVersionreplaces string-sniffing compatibility checks such asgetApiVersion().startsWith("1.")with a structured(major, minor, patch)value exposingisAtLeast(major, minor)andisCompatibleWith(other).Provenance, and detecting schema drift
TestDescriptorgained.author(...),.version(...),.tags(...)and.documentationUrl(...), plus an automatically computedschemaHash: 12 hex characters of the SHA-256 of the canonical schema — id, version, the shape of the parameters and the shape of the metrics. A backend can compare that hash across submissions and notice when a test's schema changed underneath it.Reports carry these fields (
provider,author,extensionVersion,schemaHash,documentationUrl) at the top level of each test object, and they are filtered out of themetrics{}andflags{}blocks so they stay structured metadata rather than pseudo-measurements.TestDescriptor.Builder.owner(String)is deprecated: the registry injects the owner from thePluginpassed toregister(desc, plugin). Existing call sites keep working with a deprecation warning.Events, test doubles and the rest
Two Bukkit events,
TestStartingEventandTestCompletedEvent, are fired by the host before and after each test run, with listener exceptions isolated so a misbehaving listener cannot take a run down.Metricis now a marker interface implemented byRichMetricandRichMetricSeries, andRichMetric.ofInt(long)andRichMetric.ofDouble(double)are explicit factories —ofIntdefaults to a precision of 0, so an integer counter renders as "500" and not "500.00".For unit tests,
testing.MockBenchmarkContextandtesting.CapturingCompletionlet an extension be tested without bootstrapping a Bukkit server.Two contexts that used to mislead were fixed along the way:
ctx.getLogger()actually prefixes each record with[testId]now instead of handing back the unprefixed host logger, andctx.getSender()is exposed onBenchmarkContextso an extension can route out-of-range warnings to the operator in-game rather than only to the server console.Tooling, documentation and samples
./gradlew newExtension -PextName=MyBench -PtestId=mybench.fooscaffolds a fresh extension plugin underextensions/<name>/, wired with the canonical 1.1 patterns: parameter and metric specs,TestCompletion, owner auto-injection anddepend: [VoxelBench]../gradlew checkApiImportsfails the build if any file underapi/imports outside the allowed surface — Bukkit, the JDK, and its own package. It is wired into the standardchecklifecycle, so the API stays free of internals by construction rather than by review.Five documents accompany the surface:
API_STABILITY.md(SemVer and deprecation policy, the three-tier@ApiStatusmodel, a breaking-change matrix),API_THREADING.md(which thread each method runs on, three execution patterns, the cancellation contract,cleanup()semantics, anti-patterns),API_CONVENTIONS.md(maintainer conventions),BACKEND_INGESTION_SPEC.md(an exhaustive presence matrix for every JSON key the plugin emits, sentinel-value handling and a recommended ingestion pipeline), andEXTENSION_API_REFERENCE.md, refreshed to 1.1 with a status table.The five canonical samples were migrated to the 1.1 patterns —
TestCompletion,ParamSpec,MetricSpec,ofIntand provenance metadata. One of them is new:CreateKineticBenchmark, a hybrid-only benchmark targeting the Create mod (around 100 million downloads on CurseForge), which demonstrates the probe-then-skip pattern — detect the mod's blocks throughBukkit.createBlockData(), and returnTestResult.skipped(...)when they are absent instead of failing. - 1.1.0betaarchivedMay 1, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1 · 1.1 MB · 13 ↓No longer available
Until now, VoxelBench always benchmarked in the first world your server loaded — usually the one your players live in. This release lets you pick the world instead, or have the plugin build a dedicated one for you in a single command. It also integrates with Multiverse-Core, blocks gameplay benchmarks from the console where they never worked properly, and makes sure a disconnect mid-run no longer leaves you stranded inside a bench world.
Everything here is opt-in. If you never touch
/bench world, the plugin behaves exactly like v1.0.2. There is no config change to make and no database migration to run.Pick the world your benchmarks run in
The new
/bench worldsubcommand pins a world for every future benchmark:/bench world list— list the loaded worlds, tagged[pinned],[bench]and[MV]/bench world show— show which world is currently pinned/bench world set <name>— pin a loaded world for every future run/bench world unset— clear the pin and go back to the default behaviour/bench world create <name>— create a flat, normal-environment world/bench world delete <name>— delete avoxelbench_*world
A world created this way is auto-prefixed with
voxelbench_, which is what makes deletion safe:/bench world deleteonly accepts names carrying that prefix, so the command cannot be pointed at your survival map. Its seed is fixed tobenchmark, so the terrain is the same on every server and across every run — two benchmarks measure the machine, not the landscape they happened to land in.The pin is stored in config as
benchmark.target-world, and it is honoured by every code path that touches a world:/bench start,/bench stresslimit, individual/bench testruns, the multi-run warmup, post-run region cleanup, the test lock manager — so/bench stopand the shutdown cleanup target the right world — and the confirmation message shown when a run starts. Tab completion knows the subcommand, and only suggestsvoxelbench_*names fordelete.Deletion refuses outright if a benchmark is currently running in that world, and if you delete the world that was pinned, the pin is cleared for you rather than left dangling. The user-supplied part of a world name must match
^[A-Za-z0-9_-]{1,32}$, which blocks path traversal, characters that are illegal on Windows filesystems, reserved names and lengths that overflow ext4 — and the same check runs again on delete, not only on create.Multiverse-Core, on 4.x and 5.x
The plugin detects Multiverse-Core at startup and routes world creation and deletion through it when it is present, falling back to Bukkit's native
WorldCreatorand folder removal when it is not. The feature therefore works on every server, with or without Multiverse.Detection is based on the plugin name rather than on a specific API surface, so it holds across MV major versions. Worlds you create are imported into Multiverse (
/mv import), which means they show up in/mv listand inherit your MV defaults instead of existing as something only VoxelBench knows about. Deletion observes the actual server state after each step rather than assuming a fixed sequence, so it behaves correctly whether MV 5.x's/mv removealready deleted everything or MV 4.x merely un-registered the world. Multiverse-Core is declared as asoftdependinplugin.yml, so Bukkit's loader guarantees it boots first.At startup the plugin now prints a summary of the third-party plugins it has hooked into, so you can confirm an integration was picked up instead of inferring it from behaviour.
A disconnect no longer strands you in the bench world
Disconnecting mid-benchmark — or restarting the server — used to leave you inside the benchmark world with no way back to where you were.
Your position and gamemode from before the test are now written to disk, in a per-UUID YAML file under
plugins/VoxelBench/data/player-states/. When you reconnect inside avoxelbench_*world, a listener puts you back: at the exact pre-test position when the snapshot is there, and otherwise at your bed spawn or at the spawn of the first non-benchmark world. Because the snapshot lives on disk rather than in memory, it survives a JVM crash, a server restart and a plugin reload.A SPECTATOR fix you will feel
Players teleported into a fresh Multiverse-managed benchmark world landed in SURVIVAL and immediately fell to their death. The gamemode is now re-applied three times around the teleport — before it, immediately after it, and again on the next tick — which defeats per-world gamemode plugins that were silently overriding it on entry into a new world.
Gameplay benchmarks refuse the console
Gameplay benchmarks and gameplay tests can no longer be started from the console. They always misbehaved there, silently, because they need a real player in the world; the refusal is now explicit and translated.
Hardware tests stay scriptable from the console — memory, disk, multicore, singlecore and network. The distinction is not a hard-coded list: it is keyed off each test's own
requiresPlayerPresence()metadata, so any test added later inherits the right behaviour without anyone remembering to update a list.Compatibility and polish
Minecraft's year-based version numbers are now formatted correctly: 26.1.2 displays as 26.1.2 rather than 1.1.2, and version comparisons treat a year-based major as being past every 1.x check.
The published jar carries its version in the filename (
VoxelBench-1.1.0.jar), and the version insideplugin.ymlis derived from the release tag, so the file you downloaded and the version the server reports can no longer disagree.Upgrading
Drop the new jar into
plugins/and restart. No config or database migration is needed, and every behaviour change is opt-in behind/bench world set.The Java requirement is unchanged: Java 16 or newer for Minecraft 1.17 through 1.21.x. Minecraft 26.1.x requires Java 25 LTS — that is Mojang's requirement, not ours.
- 1.0.2betaarchivedApr 27, 2026MC 1.17, 1.18, 1.19, 1.20, 1.20.6, 1.21, 26.1 · 1.1 MB · 12 ↓No longer available
VoxelBench 1.0.2 runs on Minecraft 26.1.x, and compatibility across the full 1.17 to 26.1.2 range is now checked every day by booting real Paper servers. Nothing else moves: no new features, no behaviour changes, the plugin's runtime surface is identical to 1.0.1. Drop the new jar in
plugins/, restart, and you are done.Minecraft 26.1.x is supported
Minecraft 26.1.x arrived in April 2026 with Mojang's new year-based versioning. VoxelBench is drop-in compatible with that line — 26.1, 26.1.1 and 26.1.2 — and the version is part of the compile-time CI matrix, built against the Java SE 25 LTS toolchain.
Nothing broke on the way there. There are no API breakages in 26.1.x, and the Bukkit surface VoxelBench relies on remains stable all the way back to 1.17, so supporting the new line cost nothing to the old one.
Six more entity types recognized
The entity compatibility layer gained six accessors for mobs and projectiles introduced after 1.20: armadillo (1.20.5+), breeze, bogged and wind charge (1.21+), creaking (1.21.3+) and happy ghast (1.21.5+). On a server that predates any of them, the accessor simply returns nothing — the existing supported range is unaffected.
Compatibility is now verified by booting real servers
VoxelBench is tested across the full 1.17 to 26.1.2 range, and that testing no longer stops at compilation. A new workflow boots actual Paper servers with the plugin installed — 1.17.1, 1.21 and 26.1.2 — every day at 4am UTC, on every push, and on demand. Each one is given
/bench infoand/bench tests listover stdin, and the logs are searched for the patterns that mean a plugin failed to load:Caused by:,NoSuchMethodError,ClassNotFoundException.Each version boots under the JVM its Paper bootstrap actually supports — Java 17 for 1.17.1, 21 for 1.21, 25 for 26.1.2 — because a toolchain mismatch is one of the ways this kind of check quietly stops proving anything. Server logs are kept as artifacts for seven days whether the run passed or failed, so a failure can be read after the fact rather than reproduced.
The practical consequence: a compatibility regression is caught here, before it is caught on your server.
Dependencies
The Adventure text and formatting library moves from 4.17.0 to 4.26.1, along with the MiniMessage and legacy serializer components that go with it.
Upgrading
Drop the new jar in
plugins/and restart. It is a drop-in replacement for 1.0.1: no configuration migration, no database change, nothing to adjust.Java requirement
Java 16 or newer for Minecraft 1.17 through 1.21.x. Java 25 LTS is required if you run Minecraft 26.1.x.
If you run your own backend
The obfuscated jar's hash changes with this release. That is expected — the Adventure bump shifts the bytecode — but it means the new hash has to be added to
VALID_PLUGIN_HASHESbefore reports from 1.0.2 will be accepted.If you build from source
The build now requires JDK 21; CI uses Temurin 21.
- 1.0.1alphaarchivedApr 23, 2026MC 1.21 · 1.1 MB · 18 ↓No longer available
The plugin's web dashboard was written in French: every label and message on its pages was a hardcoded French string. This release makes English the default, adds a language selector to every page, and moves the translations into files that can be extended without touching the pages.
English by default, French on request
The three dashboard pages — login, dashboard, and the API-only page — no longer carry hardcoded French strings. Their text is resolved through
DashboardI18n, a translation loader backed by.propertiesfiles. English (en_US) is now the default; French (fr_FR) is available as an option.Each of the three pages carries a language selector. The choice is stored in the
voxelbench_localecookie for one year (SameSite=Lax), so the dashboard opens in the language you picked the next time you sign in. When no cookie is present, the dashboard reads the browser'sAccept-Languageheader, and falls back toen_USwhen that yields nothing.A new language is a single file
Translations live as
.propertiesfiles underresources/dashboard/lang/. Adding a language means dropping one more file into that directory; the page code does not change.Server logs
The monitor's server logs are now written in English too, for consistency with an international audience.
- 1.0.0alphaarchivedMar 30, 2026MC 1.21 · 1.1 MB · 27 ↓No longer available
This release rebuilds the stress-limit mode, adds four tests — among them a raw single-core measurement — and lets the plugin tell you when a newer version is available.
The stress limit stops straddling the breaking point
The stress limit raises the load tier after tier until the server gives way. Each tier used to multiply the previous one's load by 1.8, so the answer came back as a wide bracket: the last tier that held, and one nearly twice as heavy that did not. The multiplier is now 1.2, which makes the ladder far finer.
Once a tier breaks, VoxelBench no longer stops there. It bisects between the broken tier and the last one that held, up to two times, to close in on the point where the server actually gives way.
Tiers the machine absorbs without effort no longer cost a full measurement: a tier still at 19.5 TPS or above after five seconds is skipped. Those five seconds are now counted on the wall clock rather than in ticks, so a falling tick rate no longer stretches them.
Base values and maximum caps were raised for every stress test type.
A stress-limit run behaves like a standard benchmark
Monitoring, the scoreboard and the metrics were aligned on the standard benchmark flow: monitoring is delegated to the shared test base class, and a single scoreboard covers the whole run. The scoreboard can also display average values, so what you read on screen matches the metric the run is actually measuring.
The run also gained the preparation a benchmark needs. The test zone is generated once and shared across tiers, the player's state is saved before the run and restored afterwards, and garbage collection is stabilised before measurement begins.
One test was measuring the wrong thing: the hopper test ran with its tick acceleration disabled in stress-limit mode. It no longer does.
Finally, every message the stress limit prints goes through the translation files, like the rest of the plugin.
Stress-limit results are readable, and can be sent
Each tier now carries a hover tooltip with its full metrics — percentiles, standard deviation, garbage collection.
A stress-limit run can also be submitted to the backend, over protocol v1.1.0.
Four new tests
- Single-core benchmark — raw single-thread operations per second, measured on the main thread.
- Projectile storm — a barrage of arrows, snowballs and eggs.
- Entity cramming — entities confined in cells, the way a mob farm confines them.
- Player world load — a multi-zone chunk-loading simulator.
The plugin tells you when it is out of date
VoxelBench now checks the backend for a newer version, asynchronously, and notifies operators when they join the server. You no longer have to go and look.