Skip to main content

Spectra Detect AMI Scanner Limitations and Troubleshooting

Limitations​

Filesystems​

The scanner mounts and walks Linux filesystems: xfs, ext4, ext3, ext2, and btrfs. Anything else is refused rather than scanned.

Encrypted and volume-managed root filesystems can't be scanned. If the largest filesystem on the asset sits inside a Logical Volume Manager (LVM) volume group or a Linux Unified Key Setup (LUKS) container, the scan fails with an explicit reason. The scanner can't activate a foreign volume group or open an encrypted volume.

note

This is a deliberate failure rather than a partial success. On such an image, the only partitions still visible are usually /boot and the EFI System Partition. Scanning those would report a successful scan having never opened the root filesystem at all.

Windows filesystems aren't supported. NTFS isn't in the list of filesystems the scanner mounts.

Instance-store volumes are out of scope. Only EBS-backed assets can be snapshotted, and only a snapshot can be scanned.

Scan Coverage​

A scan is a point-in-time picture of a stored disk. It sees what's written to the filesystem. It doesn't see running processes, memory, or anything that exists only at run time.

Filtering fails open, and so does type detection. When a file's MIME type can't be determined, the file is uploaded rather than dropped. An exclusion reduces scan volume; it isn't a guarantee that no such file is ever sent.

A capped scan covers part of the asset. scan_max_files stops files being selected once the cap is reached. The walk still runs to the end, and the remaining files are counted as skipped under the reason max_files_reached, so the file counts stay honest about how much was there. The report also carries the anomaly max_files_truncated, which forces completeness.status to partial.

The files that make the cap are the ones the walk reached first, in directory order. A low cap on a root filesystem can exhaust /bin and never reach /usr, so a clean report from a capped scan says nothing about the directories the walk never got to.

scan_max_file_size works differently. It's a per-file exclusion, like the MIME and path filters, and oversize files are counted under max_size_exceeded. It doesn't make a scan partial - a scan that skipped files only for being too large still reads complete.

Symbolic links pointing outside the image aren't scanned or reported. Following one would read a file on the scanner host rather than one in the image. They're counted as skipped.

Crash-Consistent Snapshots​

Scanning an attached EBS volume snapshots it first, so the running workload is never touched. That snapshot is crash-consistent rather than quiescent: it captures the state the filesystem would be in after a power cut, not after a clean unmount. In-flight writes may be missing, and a journal may need replay.

The report records this in metadata.source_volume_in_use. Note that attached doesn't imply running - a stopped instance's volume is still attached.

For the most consistent result on a live workload, take your own snapshot at a moment you control, quiesce the application first if it supports that, and queue the snapshot rather than the volume.

Kubernetes​

PersistentVolumes backed by EBS are discovered through their CSI tags and scanned as the underlying EBS volume. The scanner runs outside the cluster and doesn't deploy into one.

Pod names aren't available in a report. A PersistentVolume's binding to a pod happens at run time, and recovering it would require querying the Kubernetes API with additional access. Reports identify a PersistentVolume by its claim name, namespace, and cluster.

Operational Constraints​

One asset at a time per instance. Concurrency comes from running more instances, bounded by asg_max_size.

A cold start is slower. With the fleet scaled to zero, the first scan after an idle period waits for an instance to boot.

AWS Fargate can't be used. Mounting the scan volume needs host block-device access, which Fargate doesn't provide.

A same-region, same-zone copy is required. The temporary volume must be created in the availability zone of the instance that will attach it.

Filter and log-level changes apply to new instances. Settings reach an instance when it boots, so a change takes effect on instances launched afterward. Start an instance refresh to apply it to the running fleet.

Scheduled discovery doesn't deduplicate. A schedule that fires while a backlog is still draining queues those assets again, producing repeated work.

Troubleshooting​

A Scan Reports skipped​

The asset doesn't carry RLScan=true. Check the tag on the asset itself, and remember that the value is matched exactly and is case-sensitive - RLScan=True doesn't match.

aws ec2 describe-tags --filters "Name=resource-id,Values=ami-xxxxxxxxxxxxxxxxx"

A volume doesn't inherit eligibility from the instance it's attached to, and an asset tagged scanner:managed=true is never scanned regardless of what else it carries.

A skipped asset isn't a failure and writes no report. See Scan Outcomes and Exit Codes.

Nothing Is Scanned After Enabling the Event Trigger​

The trigger fires when an AMI becomes available, and the readiness event carries no tags - the tag is read later, when the scanner picks the message up. An image tagged after that point isn't triggered again. Queue it manually, or wait for the next scheduled sweep.

Instances Boot but Nothing Happens​

Check that scanner_ami_id names the scanner image. Left empty, the deployment falls back to a stock Amazon Linux image that boots and joins the group but carries no scanner. The boot script fails loudly, so the cause is visible in the instance logs.

A Scan Fails to Mount​

The scanner reports the reason it couldn't mount. Common causes:

  • The root filesystem is inside LVM or LUKS. See Filesystems. This isn't recoverable by configuration.
  • The filesystem type isn't supported. Only Linux filesystems are mounted.
  • The kernel is too old for the filesystem's features. The scanner distinguishes this from corruption by reading the kernel log, because the mount command reports a missing filesystem helper, a corrupt superblock, and a feature gap with one identical message. A feature gap means the volume is intact and readable by a newer kernel, so a newer scanner image resolves it.

A Scan Reports partial​

The scan produced a usable report, but not a whole one. Read completeness in the report:

  • files_errored above zero means some files failed analysis. Check the reason field on those lines in the verdict file.
  • An anomalies entry means something makes the results suspect rather than merely incomplete. Any anomaly forces the status to partial.
  • The anomaly max_files_truncated specifically means scan_max_files was reached. The size cap never produces it: oversize files are counted as skipped, like any other filter.

Files Come Back with No Verdict​

classification: 0 means Spectra Detect analyzed the file and returned no verdict. For text and configuration files, that's the normal result and isn't an error. A large unknown_count alongside error_count: 0 is a healthy scan, not a failing one.

classification absent is different: the file wasn't analyzed at all. Read reason and error on that line.

Analysis Requests Fail in a Hub Deployment​

The scanner refuses to poll a host it wasn't told to trust, because polling sends the API token. List the Worker hosts in reversinglabs_allowed_task_hosts, preferring a domain suffix so a pool that scales needs no redeployment. See Hub Deployments.

A plain HTTP task URL is also refused when the configured endpoint is HTTPS, rather than sending the token in clear text.

A Permission Is Granted but the Call Is Denied​

Check which role the grant is on. The instance role's permissions stop applying the moment the scanner assumes the scan role, so a grant on the instance role produces an access error even though it's plainly there. Queue and scan permissions belong on the scan role.

The scanner runs a permission check at startup for the same reason, so a missing grant surfaces before a scan begins rather than after a snapshot has been copied.

Notifications Never Arrive​

An email subscription stays pending until the recipient clicks the confirmation link. Terraform can't observe that and reports the subscription as created either way, so a silent pipeline is more often an unconfirmed subscription than a broken scanner:

aws sns list-subscriptions-by-topic --topic-arn "$(terraform output -raw sns_alert_topic_arn)"

A pending subscription shows "SubscriptionArn": "PendingConfirmation".

If the topic uses a customer-managed encryption key, publishing fails at run time without kms:GenerateDataKey and kms:Decrypt on that key - after the scan has already completed, and not at plan time.

A Report Can't Be Found in S3​

expected_report_location in a notification is computed from the configured bucket and prefix, not read back from the upload, so it can name an object that doesn't exist. A 404 there is normal and isn't proof that the report was stored.

A failed scan writes under the failures/ prefix rather than reports/, so look there.

Temporary Snapshots or Volumes Are Left Behind​

Cleanup runs in reverse order and is designed to run even when a scan is cancelled or fails. Resources that survive anyway - after an instance is terminated mid-scan, for example - are identifiable by their tag:

aws ec2 describe-volumes \
--filters "Name=tag:scanner:managed,Values=true"

aws ec2 describe-snapshots --owner-ids self \
--filters "Name=tag:scanner:managed,Values=true"

Every resource listed by those queries was created by the scanner, so it's safe to delete once no scan is running.

Finding the Logs for One Scan​

A scan that came off the queue carries a trace ID through its logs and into its report:

aws logs filter-log-events \
--log-group-name "$(terraform output -raw cloudwatch_log_group)" \
--filter-pattern '"<trace-id>"'

To see more, raise scan_log_level to debug and apply. The new level takes effect on instances launched after the change.

To inspect a host directly:

aws ssm start-session --target <instance-id>