Version 21 September 2026. One file is the only thing we need. This page explains how to produce it, and how to strip it so that even if our systems were breached, an attacker would learn nothing about yours.
The rule we hold ourselves to: we ask only for the fields that genuinely take part in the measurement. Everything else, we recommend you replace with a pseudonym. The list below is taken from the parser in the running code, not written as a general description.
.xml file.yaml or .json fileBoth produce the same result. B is the one to pick if controlling exactly what leaves your organisation matters to you.
nmap -sV -O -oX infrastructure.xml 10.0.0.0/24
-sV identifies service and version — required, because version is what vulnerability matching runs on. -O identifies the operating system. -oX writes XML — required.
Scan only ranges you are entitled to scan. The scan happens on your infrastructure, performed by you.
We read exactly these, and nothing else: the IPv4 address, the first hostname, every port whose state is open together with its protocol, service name, product and version, and the OS match — which we reduce to Windows or Linux and then discard the rest. Closed ports are ignored entirely.
hosts:
- id: app-server-01 # optional; derived from hostname if omitted
hostname: APP-01 # a pseudonym is fine
ip: 10.0.0.11 # a fake address is fine
os: Windows # Windows | Linux
osBuild: Server2019 # optional; sharpens vulnerability matching
role: server # server | workstation | database | domain-controller | ...
zone: internal # your own label for the network segment
services:
- { name: RDP, port: 3389, proto: tcp, public: false }
- { name: HTTPS, port: 443, proto: tcp, public: true, version: "nginx 1.18" }
software: # optional
- { name: OpenSSL, version: "1.0.1e" }
vulns: # optional, if you already have scan results
- { cve: CVE-2014-0160 }
This is the important part of this page.
The measurement answers: given this attack surface, which techniques would a monitoring stack configured as you declared actually catch? The attack surface is determined by operating system, open ports, services and versions. It does not depend on real hostnames or real addresses. Turning SRV-FIN-PROD-01 / 10.20.30.40 into HOST-A / 10.0.0.1 changes not one cell of the result.
hostnameipzoneos, osBuildport, protoservice.nameservice.versionroleThe same machine must carry the same pseudonym on every submission — otherwise the system reads it as a new machine and reports infrastructure drift that did not happen. The internal identifier is derived from the hostname (lowercased, non-alphanumerics become hyphens), so HOST-A becomes host-a.
Keep the pseudonym-to-real-name mapping on your side, and do not send it to us. We do not need it and do not want to hold it.
args= attribute on <nmaprun> if it contains your real ranges.What we commit to in return: if your file is missing something the measurement needs, we tell you exactly which field and which cells cannot be measured, rather than quietly guessing. See Service commitments, A1.
If I anonymise, is the report still usable? Yes. The report names the techniques your monitoring would not catch, and the content needed to close them. You map pseudonyms back to real names on your side.
We do not want to submit our whole estate. Reasonable. Start with one segment, or the group of machines that matters most. The measurement is meaningful on a subset.
Our file has thousands of hosts. Tell us beforehand. Host count drives how long the twin takes to build, and we would rather give you that number in advance than have you wait for it.