Tracepoint

Documents

Performance and Data Handling, Measured Results

Tracepoint · A&R Strategic Solutions Measurements taken 25 August 2026. All data used in these tests is synthetic. No real audit data was used or represented.


What was measured

Every number below comes from running the product's own code. The benchmark calls the same CSV parser, the same guarded engine actions, and the same export-package builder that the application uses when a person clicks Import and then Export. Nothing is stubbed, simulated, or estimated.

The benchmark runs in two places so the results can be checked two ways:

Test machine: 8-core Apple Silicon Mac, Node 22.17.0, Chromium browser engine. A government workstation of similar specification should produce comparable results. The application runs entirely on the local machine, so no network time is involved.

Each figure below is from a single recorded run, and the raw output of that run is saved in the repository. Repeating a run moves the timings by a few percent, as it would for any measurement of this kind. The results are not averaged, because averaging would hide that variance rather than disclose it.


Scale results

Each run takes a source file of audit findings, reads and maps it, creates every finding with an attested audit-trail entry, and builds the complete export package.

FindingsSource fileRead and mapCreate and attestBuild packagePackage sizeEnd to end
1000.07 MB15 ms7 ms77 ms0.66 MB0.1 sec
1,0000.73 MB91 ms61 ms343 ms5.85 MB0.5 sec
10,0007.33 MB836 ms499 ms4,005 ms57.81 MB5.3 sec

Throughput at the largest scale: 11,962 rows per second read and mapped, 20,027 findings per second created and attested.

The same runs inside the browser:

FindingsRead and mapCreate and attestBuild packagePackage sizeEnd to end
1,000102 ms105 ms719 ms5.85 MB0.9 sec
10,0001,009 ms942 ms7,971 ms57.81 MB9.9 sec

The browser is slower than the scripted harness, which is expected. Both produce the same package: the same 45 artifacts at the same sizes. The benchmark compares the artifact names and sizes in each ZIP. The two files are not byte-for-byte identical, because the package ID, the export time and the time on each trail row come from the machine's clock during the run.

For context on the scale: GAO reported 3,322 Notices of Findings and Recommendations open across the Department of Defense at the end of FY2023. The 10,000-finding test is roughly three times the entire reported DoD backlog, processed on one laptop in under six seconds.

Peak memory at 10,000 findings was 310 MB, which is within the normal working range of any modern workstation.


What the package contains

At every scale the export package contained 45 artifacts, and the validation gate confirmed the package assembled with checks passed. The five largest artifacts at 10,000 findings:

ArtifactSizeWhat it is
02-RECORDS/findings.json16.30 MBEvery finding record, complete
05-HISTORY/import-provenance.json14.11 MBEvery source row exactly as supplied, plus every conversion applied
04-REPORTS/finding-cap-report.html10.88 MBThe human-readable report
07-INTEROPERABILITY/canonical-workspace.json6.87 MBThe full workspace in open format
06-TRUST/traceseal-manifest.json2.32 MBHash manifest covering the package contents

The interoperability folder also carries an OSCAL POA&M, an eMASS POA&M CSV, and an ODCFO corrective action plan export. Their presence is verified as part of the benchmark rather than assumed.

The provenance file is large on purpose. Tracepoint keeps the original source row for every finding it creates, so a reviewer can compare what the product produced against what the source file actually said.


Handling imperfect data

Audit files arrive incomplete, inconsistent, and duplicated. A tool that quietly cleans these problems up hides them from the reviewer. Tracepoint discloses them.

A 500-row file was built with deliberate defects. The expected handling for each defect was written down before the test was run. Results:

Condition in the fileRowsWhat Tracepoint did
No title and no condition text23Refused the row and named the reason
Title blank, condition text present29Derived a title from the condition and marked it as derived
Fiscal year missing28Left it blank and flagged it as not provided
Classification the lexicon does not recognise37Applied a default and disclosed that it was defaulted
Source ID repeated inside the same file33Flagged each repeat and offered to skip it
Rows ready to create477Created only after the person confirmed

Every expectation was met exactly. Across the 477 rows the product recorded 376 individual conversions, each one retained in the export package alongside the original source value.

Three points matter more than the counts:

  1. Nothing is created until a person confirms it. The import preview shows what will be created, what will be defaulted, and what will be refused, before anything is written.
  2. A missing value is never invented. A blank fiscal year stays blank. Tracepoint does not guess a year, and no view displays a guessed year as though it were real.
  3. A default is always labelled as a default. When the product cannot recognise a classification, it applies one and says so on the record and in the export.

What the benchmark hardened

The benchmark runs on every build, and three behaviours in the current release were shaped by it.


How to read these numbers


Reproducing these results

From the application directory:

npm run bench:data      # generate the synthetic datasets
npm run bench -- 100    # or 1000, 10000, or messy

Results are written to tools/bench/result-<scale>.txt. The dataset generator uses a fixed seed, so the same files are produced on every machine. The browser version of the same benchmark is available at /bench.html?n=10000 while the development server is running.