Skip to main content

Deploying the Spectra Detect AMI Scanner

The scanner is deployed with Terraform into an account and region you already own. The templates create the queue, the scanner fleet, the storage for reports, and the permissions that hold them together. You supply the machine image, the network, and the Spectra Detect endpoint.

Before You Start​

WhatNotes
The scanner machine image IDFor your region. Supplied by ReversingLabs - see Availability.
A Spectra Detect endpoint and API tokenThe URL of a Worker or a Hub, reachable from the subnets you deploy into.
A VPC and at least one subnetExisting. The templates read them and never create them. How you lay it out is yours to decide; the scanner needs only to reach the AWS APIs and your Spectra Detect endpoint.
Terraform 1.9 or newerEarlier versions reject the variable validation the templates use.
AWS credentialsAble to create IAM roles and instance profiles, EC2 and Auto Scaling resources, SQS, S3, CloudWatch, EventBridge, and Secrets Manager resources.

What Gets Created​

The stack creates a scan queue, an Auto Scaling group of scanner instances that scales from zero, an S3 bucket for reports, a CloudWatch log group, a security group, and the roles those resources need. <env> below is the environment variable, which defaults to dev.

ResourceName
Scan queueami-scanner-<env>-scan-queue
Auto Scaling groupami-scanner-<env>-asg
S3 results bucketami-scanner-<env>-results-<account-id>
CloudWatch log group/aws/ec2/ami-scanner-<env>
Security groupami-scanner-<env>-sg
Scan runtime roleami-scanner-<env>-ecs-task-role
Instance role and profileami-scanner-<env>-ecs-instance-role, -ecs-instance-profile
API token secretami-scanner-<env>-rl-api-token-*, unless you supply your own

Optional resources are created only by the variables that enable them: an SNS topic, an EventBridge rule for AMI events, the scheduled-dispatch components, and a deploy-time role for continuous integration.

Nothing in the stack scans anything on its own. Both automatic triggers are off by default, so applying the templates never silently starts scanning.

Step 1: Set the Required Variables​

Four values have no default. Copy terraform.tfvars.example to terraform.tfvars and fill them in:

# aws/terraform.tfvars
aws_profile = "my-profile"
vpc_id = "vpc-xxxxxxxxxxxxxxxxx"
subnet_ids = ["subnet-xxxxxxxxxxxxxxxxx"]
reversinglabs_api_url = "https://spectra-detect.example.com"

scanner_ami_id = "ami-xxxxxxxxxxxxxxxxx"

terraform.tfvars is excluded from version control by the templates.

caution

Leaving scanner_ami_id empty falls back to a stock Amazon Linux image, which boots and joins the Auto Scaling group but carries no scanner. The boot script fails loudly rather than idling, so the failure is visible - but nothing will be scanned.

Step 2: Supply the API Token​

The token is passed through the environment so it stays out of the variables file:

cd aws
export TF_VAR_reversinglabs_api_token='<token>'

terraform init
terraform plan -out=tfplan
terraform apply tfplan

Terraform stores the token in a Secrets Manager secret that it creates. To reuse a secret you already own, set reversinglabs_api_token_secret_arn instead of reversinglabs_api_token.

Setting both, or neither, stops the plan before anything is created. The scanner can't authenticate without a token, and two sources of one means one of them is wrong.

Step 3: Tag the Assets You Want Scanned​

The scanner only touches assets that carry RLScan=true:

aws ec2 create-tags \
--resources ami-xxxxxxxxxxxxxxxxx \
--tags Key=RLScan,Value=true

The tag must be on the asset itself - the AMI's own tags, the volume's own tags, the snapshot's own tags. A volume doesn't inherit eligibility from the instance it's attached to.

The key and value are fixed in the scanner and can't be changed by any variable or setting. The value is matched exactly and is case-sensitive, so RLScan=True doesn't match. An asset tagged with any other value, such as RLScan=false, is an explicit opt-out, and leaves a record of the decision.

An untagged asset is skipped, not failed. The scanner reports the outcome as skipped without creating a snapshot or attaching a volume.

caution

Never apply the scanner:managed=true tag to your own assets. That tag marks the temporary snapshots and volumes the scanner creates for itself, and an asset carrying it is never scanned. It's also what the permission policy uses to decide which resources the scanner is allowed to delete.

Step 4: Run the First Scan​

Place a message on the scan queue:

TRACE_ID=$(uuidgen)
aws sqs send-message \
--queue-url "$(terraform output -raw scan_queue_url)" \
--message-body "{\"type\":\"ami\",\"id\":\"ami-xxxxxxxxxxxxxxxxx\",\"trace_id\":\"$TRACE_ID\"}"

type is ami, ec2-vol, or snapshot. trace_id must be a version 4 UUID, and it identifies the work item in the logs and in the report.

The fleet scales from zero on queue backlog, so the first scan in an idle account waits for an instance to boot - several minutes longer than later ones.

Follow the scan, then collect the report:

aws logs tail "$(terraform output -raw cloudwatch_log_group)" --follow \
--filter-pattern "\"$TRACE_ID\""

aws s3 ls "s3://$(terraform output -raw s3_results_bucket)/reports/" --recursive

The scanner logs a structured asset scan finished line carrying the outcome and the exit code it corresponds to. See Scan Outcomes and Exit Codes.

Step 5: Turn On Automatic Scanning​

Once a manual scan works, enable either trigger, or both. They place messages on the same queue and are drained by the same fleet.

enable_eventbridge_trigger = true # scan each AMI as it becomes available
enable_scan_schedule = true # sweep tagged assets on a schedule

Event-Driven Scans​

enable_eventbridge_trigger creates a rule that queues a scan whenever an AMI becomes available in the account. Scanning stays opt-in per image, because the scanner checks the tag when it picks the message up.

Tag the image before or during the build. The readiness event carries no tags, so the tag is read later, when the scanner starts work on the message. An AMI tagged after that point isn't triggered again - queue it manually, or wait for the next scheduled sweep.

Every new AMI in the account produces one queue message, tagged or not. Untagged ones are skipped in seconds, but a burst of AMI creations can briefly grow the fleet: the backlog metric counts messages, not scannable assets.

Scheduled Scans​

enable_scan_schedule turns on periodic discovery:

  1. An EventBridge Scheduler rule fires on scan_schedule_expression, which defaults to rate(1 day).
  2. The rule launches a short-lived EC2 instance that discovers tagged assets, places one message per asset on the queue, and terminates.
  3. The scanner fleet drains the queue.
caution

Set scan_schedule_expression longer than a full sweep takes. Discovery doesn't know what's already queued, so a tick that fires while a backlog is still draining queues those assets again. The scanner scans each message it receives, so the effect is repeated work rather than a stuck queue. A daily schedule is comfortable for hundreds of assets; shorten it only if a full pass reliably completes well inside the interval.

Capacity and Scaling​

VariableDefaultPurpose
instance_typet3.mediumScanner instance size. Raise it for large images - the scan volume is mounted on the instance, and extraction is processor-bound.
asg_min_size0Scale to zero when idle.
asg_max_size2Ceiling on concurrent scan hosts.
asg_desired_capacity1Instances at apply time.

The group scales on backlog per instance - outstanding messages divided by instances in service - tracking a target of 1.0, or one queued asset per instance. The scanner handles one asset at a time, which is what makes 1.0 the break-even point.

Instances require Instance Metadata Service Version 2 (IMDSv2), and carry the AWS managed AmazonSSMManagedInstanceCore policy, so you can open a shell on a scan host through AWS Systems Manager without exposing SSH:

aws ssm start-session --target <instance-id>

Monitoring​

aws logs tail "$(terraform output -raw cloudwatch_log_group)" --follow

Filter a single scan by its trace ID:

aws logs filter-log-events \
--log-group-name "$(terraform output -raw cloudwatch_log_group)" \
--filter-pattern '"<trace-id>"'

Two alarms ship with the stack. One fires when work is outstanding but nothing is being processed, which catches a wedged or idle fleet. The other fires when instances in service are sending no logs at all, which the first alarm can't see: log delivery can fail while scans still run and still write reports, so throughput looks healthy while every scan is unobservable.

To change how much the scanner logs, set scan_log_level to debug, info, warn, or error and apply. Instances read their configuration at boot, so the new level applies to instances launched after the change.

Costs​

What the deployment costs depends on your region, instance type, and scan volume. The billable resources are:

ResourceWhen it exists
EC2 instancesWhile draining the queue. The group scales from zero, so an idle deployment runs none.
EC2 instance for scheduled dispatchSeconds per scheduled run, then it terminates. Only with enable_scan_schedule.
EBS snapshot copy and volumePer scan. Both are deleted when the scan ends.
S3Reports and access logs, retained until you remove them.
CloudWatch LogsScanner output, retained 30 days by default.
SQS, Secrets Manager, SNS, EventBridgePer queue message, stored secret, and published notification.

Data leaves the VPC when the scanner submits files to Spectra Detect. The scan mode and file filters are what control that volume - see Configuration.

Removing the Deployment​

Empty the results bucket first, because S3 refuses to delete a bucket that still holds objects. Snapshots and volumes the scanner created are deleted as each scan ends, so a destroy after a completed scan leaves nothing behind.

terraform destroy