Back to all posts

Check game performance in CI before merging a change

A refactor can pass its functional tests and still add work to every frame. That holds for code written by a developer, Claude Code, or Codex: tests that check behavior don’t measure the cost of running it on the target hardware.

A performance check needs a repeatable workload and a measured baseline. Framedash’s perf-diff compares telemetry from two builds; run-profile-test launches your test command and runs that comparison afterward. Your own PC or CI runner executes the game. Before using the result to block a merge, establish what the comparison measures and how much an unchanged build varies.

Start with an unchanged baseline repeat

Choose one scenario you can run the same way each time, such as a fixed route through a level. Run the baseline, repeat it without a code change, and then run the candidate. Keep these conditions fixed and record them with the results:

  • Hardware, OS, graphics driver, and power settings
  • Engine and SDK versions, build configuration, and graphics API
  • Resolution, quality settings, VSync, and frame-rate cap
  • Scenario, camera path, warm-up, measurement duration, and telemetry settings

The repeat reveals variation that a single baseline cannot show. If unchanged runs differ enough to trigger the proposed threshold, investigate the environment or test procedure before enabling CI failure. A percentage copied from an example is not a noise estimate for your game.

Keep the comparison’s data population separate, too. perf-diff groups matching telemetry by build_id within its time window, which defaults to 30 days. Runs that reuse an ID contribute to the same aggregate; ci.scenario is metadata, not an automatic scenario filter for this command. Use distinct measurement build IDs when you need to isolate a run or configuration, and preserve the source commit in ci.commit. Both baseline and candidate must have retained data in the selected window.

Know what the P50 gate decides

perf-diff reports P50 and P95 for collected metrics. P50 is the median; the regression gate uses the change in P50. P95 describes the upper tail of the sampled values and does not control the exit code.

For an installed and configured CLI, this example gates on frame time:

framedash perf-diff \
  --baseline "$BASELINE_BUILD_ID" --candidate "$CANDIDATE_BUILD_ID" \
  --api-key-file ci-read.key --metric frame_time \
  --threshold 5 --fail-on-regression

Set both IDs to the values actually sent by the SDK, and set FRAMEDASH_PROJECT_ID to the project being measured. The private ci-read.key file contains an analytics:read key. Here, 5% is an example tolerance, not a recommended threshold.

For a positive baseline, a frame-time P50 increase from 10 ms to 10.5 ms is exactly 5% and passes this gate; 10.6 ms is a 6% increase and fails. These are illustrative values. The report’s isRegression flag can be true for an increase within this tolerance; the CLI applies --threshold separately. A P95 increase alone can still pass, so review the tail when intermittent stalls matter.

The flags and failure cases have a few consequences:

  • Without --fail-on-regression, the command reports the comparison without failing for a regression. With it, an increase strictly greater than the threshold exits 1. The default threshold is 0%.
  • --metric selects the metric used for the gate; the report can still contain other metrics. Without it, the gate evaluates all comparable metrics. An unavailable metric is skipped, so a passing frame-time result does not establish that GPU time was checked.
  • If no metric in the selected scope is comparable, the gate exits 1. That is missing evidence, not a measured slowdown. Authentication errors, invalid IDs, and other command failures can also make the job fail; read the diagnostic before changing code.

The other supported metrics are memory, gpu_time, io.read_bytes, io.read_time_ms, io.read_ops, load_time_ms, and mem.vram. Each requires the corresponding data to be collected; I/O and map-load measurements need their SDK setup. The gate treats increases as worse for all of them. For I/O and load time, a valid zero baseline moving to a positive value fails regardless of the percentage threshold, because percentage change from zero is undefined.

--map and --platform restrict the data, but don’t establish that two machines or scenarios are comparable. --metric load_time_ms cannot be combined with --map: map-load events have no map_id for that filter.

Run the scenario from CI

Once the report is useful on repeated runs, put the launch and comparison in the same job. The following is a Bash example; ./ci/profile-game.sh stands for a script you provide, not a file shipped with Framedash. It must launch the intended build, execute the scenario, finish sending telemetry, and exit with the test’s status.

framedash run-profile-test \
  --command "./ci/profile-game.sh" \
  --build-id "$CANDIDATE_BUILD_ID" --commit "$GITHUB_SHA" \
  --scenario plaza-loop --api-key-file ci-read.key \
  --baseline "$BASELINE_BUILD_ID" --metric frame_time \
  --threshold 5 --fail-on-regression

GITHUB_SHA is the commit supplied by GitHub Actions; use your CI system’s equivalent elsewhere. The baseline must already have been measured. To observe results while calibrating the check, omit --fail-on-regression.

run-profile-test supplies the child process with FRAMEDASH_BUILD_ID, FRAMEDASH_GIT_BRANCH, FRAMEDASH_GIT_COMMIT, and FRAMEDASH_TEST_SCENARIO when those values are available. After initializing the SDK, the Unity and Godot C# test entry point can pick them up with:

Framedash.TelemetrySDK.Instance.BeginAutomatedSessionFromEnvironment();
// Run the scenario, then flush and wait using your SDK's CI procedure.
Framedash.TelemetrySDK.Instance.EndAutomatedSession();

This is a session-tagging excerpt, not a complete test harness. UE5 exposes the corresponding methods on UFramedashSubsystem. The session sets the top-level build_id and the ci.branch, ci.commit, and ci.scenario attributes. The CI-integrated profiling guide covers session and runner setup. Ending the session does not prove that buffered events reached the server; follow your SDK’s engine-specific flush and wait procedure before letting the game exit.

Use separate keys for the two processes. The CLI reads telemetry with the analytics:read key in ci-read.key; the game sends it with an events:write key. Passing the read key through --api-key-file leaves FRAMEDASH_API_KEY available for game initialization to read the ingest key. Keep both keys in CI secrets or private files, outside source control and logs.

Render when the test is about rendering

UE5’s -nullrhi disables normal rendering. It can suit headless logic tests, but it does not reproduce the rendered workload or provide a meaningful GPU comparison. For graphics performance, run with rendering enabled on the intended GPU and keep those settings identical on both sides. The same caution applies to no-graphics modes in other engines.

Check that the measurement is complete

run-profile-test waits for the candidate’s performance-event count to increase from its pre-run value. That avoids accepting only old data on a rerun, but it is not proof that every event from the new run has arrived. The runner also rejects a failed or timed-out game process before attempting the regression gate.

The existing perf-diff gate has no statistical minimum sample count and does not validate all the measurement conditions above. Check the sample and session counts, scenario completion, and SDK logs as part of your test procedure. A few available samples can produce a comparison without being enough evidence for a release decision.

Its percentiles describe ingested telemetry samples, not necessarily every rendered frame. Periodic heartbeats and event snapshots can miss brief stalls, and changing event frequency can change the sample population. For Unity projects that need bounded per-frame evidence, the separate Unity performance-run pilot compares a baseline, unchanged repeat, and candidate with run-diff. It reports frame-interval distributions and completeness checks; it does not apply this P50 regression gate or add a P95 failure rule to perf-diff.

Investigate the difference, then rerun the fix

When a comparable result exceeds the threshold, first rerun under the same conditions. If the change persists, narrow the affected workload and capture it in the engine profiler. A higher frame-time or memory value identifies a symptom; it does not identify the responsible function or allocation.

If the scenario sends position-qualified events with a registered map ID, a performance heatmap can help locate expensive areas. Check the build and other active filters before attributing a hotspot to the candidate. A map is optional for the build comparison itself.

An agent can help read that evidence through Framedash’s read-only MCP tools. Aggregate tools such as get_heatmap require analytics:read; raw SQL through query additionally requires data:admin. Grant that broader scope only when the investigation needs raw queries. Read-only telemetry access does not verify the agent’s explanation or proposed code change: use the profiler to test the hypothesis, review the patch, and measure the fixed build with the same scenario.

The CLI reference covers command options. For Claude Code, the plugin bundles MCP setup and integration guidance:

claude plugin marketplace add crane-valley/framedash-claude-plugin
claude plugin install framedash@framedash

Begin with one scenario whose unchanged runs you understand. Save its baseline, candidate report, test settings, and the later verification result together in the pull request. That gives the reviewer enough context to decide whether the performance change is acceptable.

Compare a baseline and candidate on your own runner SDKs are available for Unity, UE5, and Godot (C#).
Start for free UnityUE5Godot (C#)