Skip to main content

Spectra Detect AMI Scanner Results and Notifications

Each scan produces two artifacts, and optionally a notification. All three are stable contracts you can integrate against: fields shared between them use the same names and the same types, and both documents carry the same schema_version.

ArtifactContentSize
<timestamp>.jsonSummary: counts, completeness, provenance, and the malicious and suspicious findingsFixed, around 1 KB
<timestamp>.verdicts.ndjson.gzOne line per analyzed file - the full per-file recordScales with file count

The summary deliberately doesn't list every file. It stays a fixed-size document regardless of how large the asset is, and the compressed newline-delimited JSON (NDJSON) file is the sole full record. A consumer that needs per-file detail reads the NDJSON file.

Where Results Are Written​

When S3 output is enabled, reports are written under the reports prefix of the results bucket:

s3://<bucket>/reports/<resource-id>/<YYYYMMDD>T<HHMMSS>Z.json

A failed scan writes to a failures/ prefix instead, using a record that has no classification counts at all. malicious_count: 0 is therefore unrepresentable on a failed scan, rather than merely suppressed - nothing in that document can be misread as "the scan found nothing".

Summary Report​

{
"schema_version": "2.2.0",
"generated_at": "2026-09-04T08:39:17.245088911Z",
"result": {
"resource_id": "ami-xxxxxxxxxxxxxxxxx",
"region": "us-east-1",
"account_id": "123456789012",
"start_time": "2026-09-04T08:16:15.917621723Z",
"end_time": "2026-09-04T08:39:17.245088911Z",
"duration_seconds": 1381.327467199,
"duration_human": "23m1.327467199s",
"scan_mode": "full-filesystem",
"total_files": 27556,
"malicious_count": 0,
"suspicious_count": 0,
"clean_count": 6815,
"unknown_count": 20741,
"error_count": 0
},
"summary": {
"total_size": 0,
"malicious_files": [],
"suspicious_files": []
},
"metadata": {
"account_id": "123456789012",
"availability_zone": "us-east-1a",
"filesystem_type": "xfs",
"instance_id": "i-xxxxxxxxxxxxxxxxx",
"mount_root": "/mnt/ami-scan/snap-xxxxxxxxxxxxxxxxx-db1175ae77f0439796f5acf95d160d84",
"mounted_device": "/dev/nvme1n1p4",
"region": "us-east-1",
"resource_id": "ami-xxxxxxxxxxxxxxxxx",
"scan_mode": "full-filesystem",
"scanner_version": "49593ab",
"snapshot_id": "snap-xxxxxxxxxxxxxxxxx",
"volume_id": "vol-xxxxxxxxxxxxxxxxx"
},
"completeness": {
"status": "complete",
"files_discovered": 32387,
"files_analyzed": 27556,
"files_errored": 0,
"files_skipped": 4831
}
}

Top-Level Fields​

FieldTypeMeaning
schema_versionstringThe version of the report's document shape, not of the scanner build. Build provenance is in metadata.
generated_atRFC 3339When the report was written.

result - What Was Scanned and What Was Found​

FieldTypeMeaning
resource_idstringThe asset that was scanned: an AMI, volume, or snapshot ID. Named resource_id rather than ami_id because the scanner isn't AMI-only, and this is the key a consumer stores and correlates on.
region, account_idstringWhere the asset lives.
start_time, end_timeRFC 3339Scan boundaries, including the snapshot copy and the mount.
duration_secondsfloatElapsed seconds.
duration_humanstringThe same value in a readable form.
scan_modestringWhich filter set ran.
total_filesintFiles analyzed, not files discovered. Equal to completeness.files_analyzed.
malicious_count, suspicious_count, clean_count, unknown_countintThe verdict tally. clean_count counts goodware.
error_countintFiles that produced no verdict. These are counted here, and are never folded into unknown_count.

summary - The Actionable Subset​

FieldTypeMeaning
total_sizeintNot populated by the scan path; currently always 0.
malicious_filesarrayFull records for malicious findings.
suspicious_filesarrayThe same, for suspicious findings.

An empty malicious_files is an empty array, never null and never omitted. In a security report, null invites reading as "unknown" rather than "none found", and an empty finding list is itself a finding.

metadata - Provenance​

Keys are omitted when the value is unavailable, rather than emitted empty.

FieldMeaning
instance_idWhich scanner host ran the scan.
instance_image_id, availability_zoneThe image the scanner host itself booted from, and the zone it ran in. This is the scanner's own runtime image, not the image under scan - that's resource_id. A value that fails validation is omitted rather than published as a plausible-looking ID that names nothing.
trace_id, attempt_idtrace_id identifies the work item, and is assigned when the item is queued, so a redelivered message keeps the same one. It's the key that joins a report back to the request that caused it. attempt_id identifies one execution of that work item. Both are needed, because two reports for the same asset are otherwise indistinguishable, and a retry can't be told apart from its original.
asset_typeWhat kind of thing resource_id names: ami, ec2-vol, or snapshot. Without it the ID alone is ambiguous, because a snap- ID could be a snapshot someone asked to scan or the temporary copy the scanner took of an image.
source_volume_in_usePresent and "true" only when the scanned EBS volume was attached when its snapshot was taken. See Crash-Consistent Snapshots.
snapshot_id, volume_idThe temporary resources the scanner created and cleaned up.
mounted_device, filesystem_typeWhat was actually mounted, after device resolution.
mount_rootThe directory on the scanner host where the scanned filesystem was mounted. Every path in the report and in the verdict records is relative to it, so /etc/passwd in a finding means <mount_root>/etc/passwd - never the scanner host's own /etc/passwd.
scanner_version, scanner_commit, scanner_build_dateWhich scanner build produced the report. Omitted entirely unless the build system stamped the binary. An absent key means "provenance unknown"; the placeholder dev is never emitted as though it were a real version.
failure_reasonPresent only on a failed scan. States why it failed.

completeness - Whether the Result Is Whole​

FieldTypeMeaning
statusstringcomplete, partial, or failed.
files_discoveredintFiles the walk found.
files_analyzedintFiles that produced a verdict.
files_erroredintFiles that failed analysis.
files_skippedintFiles excluded by the mode's filters. In the example above, 32,387 − 27,556 = 4,831.
reasonsobjectA tally of why files were skipped or failed, keyed by reason - for example mime_category_excluded or symlink_escapes_root. Present only when there is something to count.
anomaliesarrayPresent only when something makes the results suspect rather than merely incomplete - for example max_files_truncated, recorded when scan_max_files was reached. Any anomaly forces status: partial.

A partial scan still writes a usable report. The status exists so a partial result is distinguishable from a total failure, rather than being discarded.

Scan Outcomes and Exit Codes​

Every scan ends in one of four outcomes, reported as the outcome field of the per-asset log line. Each has a matching exit code:

CodeOutcomeMeaning
0succeededThe asset was scanned and every selected file produced a verdict.
1failedThe scan couldn't produce a usable result. The report is written under failures/.
3partialThe asset was scanned, but the result isn't whole. A usable report is still written.
4skippedThe asset carries no RLScan=true tag. Nothing was scanned, and no temporary resources were created.

The outcome and completeness.status above are not the same field. The status has three values - complete, partial, failed - and describes a scan that ran, so a skipped asset has no status at all: nothing was scanned, and no report was written. A clean scan reads succeeded in the log and complete in the report.

On the deployed fleet, the log line is where an asset's outcome lives. The scanner handles many assets in one process, so its exit code describes the health of that process, not the assets it touched - a run that ends cleanly exits 0 even if every asset in it failed. The codes above describe the asset only when a single asset is scanned in the foreground.

Discovery Codes​

Two further codes report on asset discovery only, and say nothing about any scan. They come from the discovery pass that a scheduled sweep runs, which finds tagged assets and queues them - it never scans one:

CodeMeaning
5Discovery failed entirely. Nothing was queued, and nothing will be until it's fixed.
6Discovery partly failed. Work was queued, but one class of asset has stopped being discovered.
caution

A non-zero code doesn't always mean something went wrong.

4 means the asset wasn't marked for scanning. That's the normal outcome for untagged assets, not an error.

6 reports a run that covered less than the whole account. The scheduled dispatch instance records its exit code in CloudWatch Logs as dispatch run finished exit_code=6, but work was still queued, so the queue drains and the fleet looks busy. Nothing else signals it.

Per-File Verdicts​

The verdict file is gzipped NDJSON - one JSON object per line, one line per analyzed file. It's streamable, so no consumer needs to hold the whole set in memory:

zcat 20260904T083917Z.verdicts.ndjson.gz | jq -c 'select(.classification == 3)'
{"path":"/usr/bin/[","sha256":"6268a27449c0c7da05a8be6b115663e3a9fb05dd4d94fbf4cb16bbd119e39845","reported_sha256":"6268a27449c0c7da05a8be6b115663e3a9fb05dd4d94fbf4cb16bbd119e39845","md5":"f92a022fcd8f9e2de29e65c1f51b1c52","sha1":"824b8606af2b74eabe90c1b9329965b3fb0f5316","rha0":"824b8606af2b74eabe90c1b9329965b3fb0f5316","size":52864,"classification":0,"risk_score":0,"threat_name":null,"file_type":"ELF64 Little","file_link":null,"propagated":false,"task_id":105770,"task_url":"https://worker.example.com/api/tiscale/v1/task/105770","processed_at":"2026-09-04T08:24:14Z","error":null,"reason":null}
FieldTypeMeaning
pathstringAbsolute path inside the scanned filesystem, not on the scanner host. Resolve it against metadata.mount_root to get the path the scanner read.
sizeintFile size in bytes.
sha256stringComputed locally, over the bytes that were uploaded.
reported_sha256string or nullThe SHA-256 Spectra Detect recorded for the file it analyzed. Deliberately separate from sha256: one covers what was sent, the other what was received. A disagreement means the bytes analyzed weren't the bytes uploaded, and collapsing the two fields would hide exactly that.
md5, sha1, rha0string or nullHashes as reported by Spectra Detect. rha0 is the ReversingLabs functional similarity hash.
classificationint or absent0 unknown, 1 goodware, 2 suspicious, 3 malicious. Absent when the file produced no verdict at all.
risk_scoreint0 to 10. Meaningful only alongside a classification - ignore it entirely when there's no verdict.
threat_namestring or nullThe threat name, such as Win32.Trojan.Generic.
file_typestring or nullThe type Spectra Detect determined, such as Binary, Text, ELF64 Little, PE+.
file_linkstring or nullAbsolute URL to the Spectra Detect copy of the file, when one was retained.
propagatedbool or nullThe Spectra Detect propagated flag on the verdict, passed through as it is.
task_id, task_urlint or null, string or nullThe Spectra Detect task, and an absolute URL for it that can be followed directly.
processed_atRFC 3339 or nullWhen Spectra Detect finished analyzing the file.
errorstring or nullError detail for a per-file failure.
reasonstring or nullA machine-readable failure category: submit_failed, report_expired, report_pending, poll_timeout, parse_failed, cancelled, or internal_panic.

Reading a Verdict Correctly​

  • classification absent - the file wasn't analyzed. Ignore risk_score, and read reason and error instead.
  • classification: 0 - analyzed, but Spectra Detect returned no verdict. This is a normal result for text and configuration files, and isn't an error. In the example scan above, 20,741 of 27,556 files landed here with zero errors.
  • classification: 1 - goodware. A non-zero risk_score and a *.Format.Graylisting threat name are still expected here, because graylisting reflects file format rather than a threat.
  • classification: 2 or 3 - suspicious or malicious. These files also appear in the summary report's suspicious_files and malicious_files.
note

Absence is never a zero value. Optional fields are serialized as explicit null rather than being omitted or emitted as "" or 0, so a consumer can tell "absent" from "this build never emits the field". A clean file has no threat, not a threat named "". classification is the one exception: it's omitted rather than set to null, for compatibility with consumers that already treat a missing key as "no verdict".

Notifications​

When an SNS topic is configured, the scanner publishes a JSON summary at the end of every scan - not only scans that found something. A clean result publishes too, with malicious_count and suspicious_count at zero. That's deliberate: if you only ever hear from the scanner when it finds something, you can't tell a quiet week from a scanner that stopped running.

{
"schema_version": "2.2.0",
"resource_id": "ami-xxxxxxxxxxxxxxxxx",
"region": "us-east-1",
"account_id": "123456789012",
"scan_time": "2026-09-04T08:39:17Z",
"total_files": 27556,
"malicious_count": 0,
"suspicious_count": 0,
"error_count": 0,
"status": "complete",
"files_discovered": 32387,
"files_analyzed": 27556,
"files_errored": 0,
"files_skipped": 4831,
"anomaly_count": 0,
"expected_report_location": "s3://my-scan-reports-bucket/reports/ami-xxxxxxxxxxxxxxxxx/20260904T083917Z.json"
}

No numeric field is omitted when it's zero. "malicious_count": 0 is a finding, and it must not look like a field an older producer never sent. Use schema_version to tell producer generations apart - it's the only field whose absence carries meaning.

The notification carries counts and status, not file paths. To see which files were flagged, read the report from S3.

caution

expected_report_location is computed from the configured bucket and prefix, not read back from the upload, so it can name an object that doesn't exist. Treat a 404 as normal, and don't use the field as proof that the report was stored.

Alerting on Detections Only​

The scanner sets no SNS message attributes, so a subscription filter policy has to use FilterPolicyScope = "MessageBody" and match on malicious_count. Filtering in whatever consumes the topic works equally well.

Configuring the Topic​

The deployment can either create the topic or publish to one you already own. The two are mutually exclusive, and setting both stops the plan rather than picking one - either choice could be the wrong one, and the wrong one publishes to a topic nobody is subscribed to, which looks exactly like a scanner finding nothing.

# Let the deployment create and own the topic.
create_sns_topic = true
sns_subscription_protocol = "email"
sns_subscription_endpoint = "security@example.com"
# Or publish to a topic you already own. Nothing about it is managed here.
sns_alert_topic_arn = "arn:aws:sns:us-east-1:123456789012:my-topic"

A topic the deployment creates is encrypted, tagged, and carries a policy allowing publication only from principals in your account. terraform output sns_alert_topic_arn reports whichever topic is in play.

Topic Encryption​

sns_topic_kms_master_key_id defaults to the AWS managed key alias/aws/sns. Encryption isn't optional: the notification names the asset that was scanned and what was found in it, so the message body is the finding.

Supply a customer-managed key ARN if you need cross-account subscribers, which the AWS managed key doesn't permit, or your own rotation policy. When the deployment creates the topic, the scanner is then also granted kms:GenerateDataKey and kms:Decrypt on that key. A topic you bring yourself with sns_alert_topic_arn gets no such grant - attach it yourself. Without those grants, publishing fails at run time with an access error, after a scan has already completed, and not at plan time.

Joining It All Together​

For a scan that came off the queue, trace_id is the key that joins everything:

  • The queue message carries it, because the sender assigned it.
  • Every log line for the scan carries it, so aws logs filter-log-events --filter-pattern '"<trace-id>"' returns the whole scan.
  • The report's metadata.trace_id carries it, so a report in S3 can be traced back to the request that caused it.

A scan run outside the queue has no trace_id, because there's no work item behind it.