Live:CloudOps Webinars & Hands-on Workshops ·Register ↗
Skip to main content

Amazon OpenSearch Service Monitoring

Overview​

Amazon OpenSearch Service publishes cluster and node metrics to CloudWatch in the AWS/ES namespace automatically. For deeper Prometheus-compatible monitoring — including per-node JVM, indexing rates, search latencies, and shard-level detail — the AWS Managed Collector can scrape your VPC-access domain and export metrics to Amazon Managed Service for Prometheus (AMP) or CloudWatch's PromQL store.

This combination gives you both the native CloudWatch alarms path and the full PromQL query capability with Grafana dashboards, without running any self-managed Prometheus infrastructure.

OpenSearch also functions as a log analytics backend itself (ingesting logs from ECS, EKS, EC2 via Fluent Bit or Logstash), but this entry focuses on monitoring the OpenSearch service health rather than using it as a logging destination.

For complete reference material, see Monitoring Amazon OpenSearch Service.

Prerequisites​

  • Amazon OpenSearch Service domain with VPC access (public-access domains are not supported by the managed collector)
  • At least two subnets in different Availability Zones (same VPC as the domain)
  • Security group allowing the collector to reach the domain endpoint over HTTPS (port 443)
  • IAM permissions: aps:CreateScraper, es:DescribeDomain, es:ESHttpGet
  • (Optional) AMP workspace for PromQL-based querying
  • (Optional) AMG workspace with AMP or CloudWatch data source configured

Architecture​

OpenSearch pipeline architecture with managed collector

┌──────────────────────────────────────────────────────────────┐
│ OpenSearch Domain (VPC access) │
│ │
│ ┌────────────┐ ┌────────────┐ ┌────────────┐ │
│ │ Data Node │ │ Data Node │ │ Data Node │ │
│ │ (metrics) │ │ (metrics) │ │ (metrics) │ │
│ └─────┬──────┘ └─────┬──────┘ └─────┬──────┘ │
│ └────────────────┼────────────────┘ │
│ │ :443 HTTPS │
└─────────────────────────┼────────────────────────────────────┘
│ Pull (private VPC)
▼
┌──────────────────────────┐
│ AWS Managed Collector │
│ (fully managed) │
└────────────┬─────────────┘
│
┌────────────┼────────────┐
▼ ▼
┌──────────────────────┐ ┌──────────────────────┐
│ Amazon Managed │ │ Amazon CloudWatch │
│ Prometheus (AMP) │ │ (PromQL store) │
└──────────┬───────────┘ └──────────────────────┘
│
▼
┌──────────────────────┐
│ Amazon Managed │
│ Grafana (AMG) │
└──────────────────────┘

Deploy​

Step 1: Configure the domain security group​

Add an inbound rule allowing HTTPS from the collector's security group:

ProtocolPortSourcePurpose
TCP443Scraper SGMetric collection endpoint

Step 2: Create scrape configuration​

Create opensearch-config.yaml:

global:
external_labels:
domain_name: my-opensearch-domain

scrape_configs:
- job_name: opensearch-exporter
scrape_interval: 60s

The managed collector resolves domain endpoints automatically — you do not specify targets manually.

Step 3: Create the managed collector scraper​

aws amp create-scraper \
--alias "opensearch-metrics-scraper" \
--source '{
"vpcConfiguration": {
"subnetIds": ["subnet-abc123", "subnet-def456"],
"securityGroupIds": ["sg-0123456789abcdef0"]
}
}' \
--exporters '[
{
"openSearchConfiguration": {
"domainArn": "arn:aws:es:us-west-2:123456789012:domain/my-opensearch-domain"
}
}
]' \
--scrape-configuration configurationBlob=$(base64 -w 0 opensearch-config.yaml) \
--destination '{
"ampConfiguration": {
"workspaceArn": "arn:aws:aps:us-west-2:123456789012:workspace/ws-abc123"
}
}'

Replace ampConfiguration with cloudWatchConfiguration (using a dataset ARN) to send metrics to CloudWatch PromQL store.

Step 4: Set up CloudWatch alarms on native metrics​

aws cloudwatch put-metric-alarm \
--alarm-name "OpenSearch-ClusterRed" \
--namespace AWS/ES \
--metric-name "ClusterStatus.red" \
--dimensions Name=DomainName,Value=my-opensearch-domain Name=ClientId,Value=123456789012 \
--statistic Maximum \
--period 60 \
--threshold 1 \
--comparison-operator GreaterThanOrEqualToThreshold \
--evaluation-periods 1 \
--alarm-actions arn:aws:sns:us-west-2:123456789012:ops-alerts

Validate​

  1. Check scraper status:

    aws amp describe-scraper --scraper-id s-1234abcd-5678-90ef

    Wait for ACTIVE status.

  2. Query cluster health (CloudWatch Query Studio or AMP):

    opensearch_cluster_health_status
  3. Query node-level JVM heap:

    opensearch_jvm_mem_heap_used_in_bytes / opensearch_jvm_mem_heap_max_in_bytes * 100
  4. Check the vended dashboard: CloudWatch → Dashboards → Templates → OpenSearch OTel.

Key metrics to monitor​

MetricMeaningAlert Threshold
ClusterStatus.red (CW)At least one primary shard unassigned= 1
FreeStorageSpace (CW)Available disk per node< 20% of total
JVMMemoryPressure (CW)JVM heap utilization> 80%
opensearch_indices_search_query_time_in_millis (Prom)Search latencyTrending upward
opensearch_cluster_health_number_of_nodes (Prom)Node count< expected

Troubleshoot​

SymptomLikely CauseFix
Scraper stuck in CREATINGSecurity group blocks HTTPS from collector to domainAdd inbound rule for port 443 from scraper SG to domain SG
No Prometheus metrics, CW metrics work fineDomain uses public access (unsupported)Migrate domain to VPC access or rely on native CW metrics only
ClusterStatus.red alarm firingPrimary shard unassigned (node failure, disk full)Check FreeStorageSpace; increase instance count or storage; see Red cluster troubleshooting
JVM memory pressure > 90%Field data or aggregation cache pressureReduce field data usage; increase instance size; enable circuit breakers
Indexing latency increasingBulk queue full or merge pressureCheck ThreadpoolWriteQueue; reduce bulk request size; add data nodes