Every API call made to AWS is an event. An engineer runs terraform apply and creates three resources. A Lambda function assumes a role and reads from S3. A compromised access key calls DescribeInstances from an IP in a country your team has never worked from. CloudTrail captures all of this by default for management events, and with data events enabled it captures object-level S3 and Lambda invocation activity too.
Without CloudTrail, security investigations become guesswork. Compliance audits require screenshots rather than structured evidence. Root cause analysis for configuration changes that broke something turns into archaeology. In this article I will walk through setting up CloudTrail at the organization level, archiving logs durably to S3, validating log integrity, and the CloudWatch Insights queries that surface useful information during an incident.
Organization-Level Trail
Create a single trail in the management account with is_organization_trail = true. This captures events from every account in the organization and delivers them to a central S3 bucket. One trail, one bucket, complete coverage across every account.
resource "aws_s3_bucket" "cloudtrail" {
bucket = "org-cloudtrail-logs-${data.aws_caller_identity.current.account_id}"
}
resource "aws_s3_bucket_versioning" "cloudtrail" {
bucket = aws_s3_bucket.cloudtrail.id
versioning_configuration { status = "Enabled" }
}
resource "aws_s3_bucket_server_side_encryption_configuration" "cloudtrail" {
bucket = aws_s3_bucket.cloudtrail.id
rule {
apply_server_side_encryption_by_default {
sse_algorithm = "aws:kms"
kms_master_key_id = aws_kms_key.cloudtrail.arn
}
bucket_key_enabled = true
}
}
resource "aws_s3_bucket_public_access_block" "cloudtrail" {
bucket = aws_s3_bucket.cloudtrail.id
block_public_acls = true
block_public_policy = true
ignore_public_acls = true
restrict_public_buckets = true
}
resource "aws_s3_bucket_lifecycle_configuration" "cloudtrail" {
bucket = aws_s3_bucket.cloudtrail.id
rule {
id = "transition-to-ia"
status = "Enabled"
transition {
days = 90
storage_class = "STANDARD_IA"
}
transition {
days = 365
storage_class = "GLACIER"
}
expiration {
days = 2555
}
}
}
resource "aws_s3_bucket_policy" "cloudtrail" {
bucket = aws_s3_bucket.cloudtrail.id
policy = jsonencode({
Version = "2012-10-17"
Statement = [
{
Sid = "AWSCloudTrailAclCheck"
Effect = "Allow"
Principal = { Service = "cloudtrail.amazonaws.com" }
Action = "s3:GetBucketAcl"
Resource = aws_s3_bucket.cloudtrail.arn
},
{
Sid = "AWSCloudTrailWrite"
Effect = "Allow"
Principal = { Service = "cloudtrail.amazonaws.com" }
Action = "s3:PutObject"
Resource = "${aws_s3_bucket.cloudtrail.arn}/AWSLogs/*"
Condition = {
StringEquals = {
"s3:x-amz-acl" = "bucket-owner-full-control"
}
}
}
]
})
}
resource "aws_cloudwatch_log_group" "cloudtrail" {
name = "/aws/cloudtrail/organization"
retention_in_days = 90
kms_key_id = aws_kms_key.cloudtrail.arn
}
resource "aws_cloudtrail" "organization" {
name = "organization-trail"
s3_bucket_name = aws_s3_bucket.cloudtrail.bucket
is_multi_region_trail = true
is_organization_trail = true
include_global_service_events = true
enable_log_file_validation = true
cloud_watch_logs_group_arn = "${aws_cloudwatch_log_group.cloudtrail.arn}:*"
cloud_watch_logs_role_arn = aws_iam_role.cloudtrail_cloudwatch.arn
kms_key_id = aws_kms_key.cloudtrail.arn
event_selector {
read_write_type = "All"
include_management_events = true
data_resource {
type = "AWS::S3::Object"
values = ["arn:aws:s3:::"]
}
data_resource {
type = "AWS::Lambda::Function"
values = ["arn:aws:lambda"]
}
}
insight_selector {
insight_type = "ApiCallRateInsight"
}
insight_selector {
insight_type = "ApiErrorRateInsight"
}
tags = {
Environment = "production"
ManagedBy = "terraform"
}
}
enable_log_file_validation = true makes CloudTrail generate a digest file every hour that contains a hash of every log file delivered in that hour. If a log file is tampered with or deleted, the digest file does not match and you know your audit trail has been altered. This is a hard requirement for SOC 2 and many other compliance frameworks.
CloudTrail Insights detects unusual API call rates and error rates compared to historical baselines. If an API like RunInstances or DeleteSecurityGroup is called at 10 times its normal rate, Insights generates a finding. Enable both ApiCallRateInsight and ApiErrorRateInsight. The error rate insight catches credential stuffing and failed enumeration attempts that the call rate insight misses.
S3 Data Events: What to Enable and What Not To
Enabling S3 data events at “arn:aws:s3:::” captures every GetObject, PutObject, and DeleteObject across every bucket in the account. For an account with high S3 traffic this can be expensive. CloudTrail charges $0.10 per 100,000 data events. An account doing 100 million S3 operations per day accumulates $100 in CloudTrail charges per day from data events alone.
The pragmatic approach is to enable data events only for buckets that hold sensitive data: buckets tagged Sensitivity=high, buckets containing customer PII, buckets used for backup storage. Scope the data resource ARN to specific buckets rather than the wildcard.
CloudWatch Insights Queries for Incident Response
Save these queries before you need them. Running them under pressure during an incident is much slower than opening a saved query.
Who has been using a specific access key:
fields @timestamp, eventName, sourceIPAddress, userAgent, errorCode
| filter userIdentity.accessKeyId = "AKIAIOSFODNN7EXAMPLE"
| sort @timestamp desc
| limit 200
All console sign-ins in the last 24 hours:
fields @timestamp, userIdentity.userName, sourceIPAddress, userAgent, responseElements.ConsoleLogin
| filter eventName = "ConsoleLogin"
| sort @timestamp desc
| limit 100
IAM changes in the last 7 days:
fields @timestamp, userIdentity.arn, eventName, requestParameters
| filter eventSource = "iam.amazonaws.com"
| filter eventName in [
"CreateUser", "DeleteUser", "AttachUserPolicy", "DetachUserPolicy",
"CreateAccessKey", "DeleteAccessKey",
"CreateRole", "DeleteRole", "AttachRolePolicy", "DetachRolePolicy",
"PutRolePolicy", "DeleteRolePolicy"
]
| sort @timestamp desc
| limit 200
Denied API calls, useful for detecting credential probing:
fields @timestamp, userIdentity.arn, eventName, sourceIPAddress, errorCode, errorMessage
| filter errorCode in ["AccessDenied", "UnauthorizedOperation"]
| stats count() as denied_count by userIdentity.arn, sourceIPAddress
| sort denied_count desc
| limit 50
CloudTrail Alarms for Critical Events
resource "aws_cloudwatch_log_metric_filter" "root_login" {
name = "root-account-login"
pattern = "{ $.userIdentity.type = \"Root\" && $.eventType != \"AwsServiceEvent\" }"
log_group_name = aws_cloudwatch_log_group.cloudtrail.name
metric_transformation {
name = "RootAccountLogin"
namespace = "SecurityMetrics"
value = "1"
}
}
resource "aws_cloudwatch_metric_alarm" "root_login" {
alarm_name = "root-account-login"
comparison_operator = "GreaterThanThreshold"
evaluation_periods = 1
metric_name = "RootAccountLogin"
namespace = "SecurityMetrics"
period = 60
statistic = "Sum"
threshold = 0
alarm_description = "Root account has been used to log in"
alarm_actions = [aws_sns_topic.security_critical.arn]
}
resource "aws_cloudwatch_log_metric_filter" "unauthorized_api" {
name = "unauthorized-api-calls"
pattern = "{ ($.errorCode = \"*UnauthorizedAccess*\") || ($.errorCode = \"AccessDenied\") }"
log_group_name = aws_cloudwatch_log_group.cloudtrail.name
metric_transformation {
name = "UnauthorizedAPICalls"
namespace = "SecurityMetrics"
value = "1"
}
}
resource "aws_cloudwatch_metric_alarm" "unauthorized_api" {
alarm_name = "unauthorized-api-calls-high"
comparison_operator = "GreaterThanThreshold"
evaluation_periods = 1
metric_name = "UnauthorizedAPICalls"
namespace = "SecurityMetrics"
period = 300
statistic = "Sum"
threshold = 20
alarm_description = "High number of unauthorized API calls in 5 minutes"
alarm_actions = [aws_sns_topic.security_alerts.arn]
}
The root account login alarm fires on any use of the root account, threshold zero. Root should never be used for day-to-day operations. Any use is worth immediate investigation. The unauthorized API alarm fires when more than 20 authorization failures occur in 5 minutes, which is a signal of credential probing or a misconfigured application.
Closing Thoughts
CloudTrail is not optional. Every production AWS account needs it enabled, every region covered, logs delivered to a central secure bucket with integrity validation on. The cost is minimal relative to the value during the incidents and audits that will eventually require it.
Set it up at the organization level once. Enable log file validation. Enable CloudWatch Logs delivery for real-time alerting on critical events. Save your incident response queries before you need them. Then treat CloudTrail as the source of truth for every question about what happened in your AWS environment and when.
Enjoy the cloud.
Osama
#AWS #CloudTrail #AuditLogging #CloudSecurity #Compliance #Terraform #InfrastructureAsCode #AmazonWebServices #SolutionsArchitect #CloudComputing #SecurityEngineering #IncidentResponse #DevSecOps #TechBlog #CloudInfrastructure #CloudGovernance #CloudWatch #S3 #IAM #ForensicsReady
Leave a comment