EKS Observability : அத்தியாவசிய மெட்ரிக்குகள்
தற்போதைய நிலை
Monitoring என்பது infrastructure மற்றும் application உரிமையாளர்கள் தங்கள் systems-இன் historical மற்றும் current state-ஐ பார்க்கவும் புரிந்துகொள்ளவும் உதவும் தீர்வாகும், வரையறுக்கப்பட்ட மெட்ரிக்குகள் அல்லது logs சேகரிப்பதில் கவனம் செலுத்துகிறது.
Monitoring பல ஆண்டுகளாக பரிணாம வளர்ச்சி அடைந்துள்ளது. சிக்கல்களை debug மற்றும் troubleshoot செய்ய debug மற்றும் dump logs பயன்படுத்துவதில் தொடங்கி, syslogs, top போன்ற command-line tools பயன்படுத்தி basic monitoring-க்கு முன்னேறி, dashboard-இல் visualize செய்யக்கூடியதாக மாறியுள்ளது. Cloud-இன் வருகையும் scale-இன் அதிகரிப்பும், இன்று நாம் முன் எப்போதும் இல்லாத அளவுக்கு அதிகமாக tracking செய்கிறோம். Industry Observability-க்கு அதிகமாக மாறியுள்ளது, இது infrastructure மற்றும் application உ ரிமையாளர்கள் தங்கள் systems-ஐ actively troubleshoot மற்றும் debug செய்ய உதவும் தீர்வாகும். Observability மெட்ரிக்குகளிலிருந்து derived patterns-ஐ பார்ப்பதில் அதிகம் கவனம் செலுத்துகிறது.
மெட்ரிக்குகள், ஏன் முக்கியம்?
மெட்ரிக்குகள் என்பது உருவாக்கப்பட்ட நேரத்தின் வரிசையில் வைக்கப்படும் numerical values தொடராகும். உங்கள் environment-இல் உள்ள servers எண்ணிக்கை, disk usage, அவை handle செய்யும் requests per second, அல்லது இந்த requests-ஐ complete செய்வதில் உள்ள latency வரை அனைத்தையும் track செய்ய பயன்படுகின்றன. மெட்ரிக்குகள் உங்கள் systems எவ்வாறு perform செய்கின்றன என்பதை சொல்லும் data ஆகும். சிறிய அல்லது பெரிய கிளஸ்டர் இயக்கினாலும், systems health மற்றும் performance பற்றிய insights பெறுவது, improvement areas-ஐ identify செய்யவும், issue-ஐ troubleshoot மற்றும் trace செய்யவும், workloads performance மற்றும் efficiency-ஐ ஒட்டுமொத்தமாக மேம்படுத்தவும் உதவுகிறது. இந்த மாற்றங்கள் உங்கள் கிளஸ்டரில் செலவிடும் நேரம் மற்றும் resources-ஐ பாதிக்கலாம், இது நேரடியாக cost-ஆக மாறுகிறது.
மெட்ரிக்குகள் சேகரிப்பு
EKS கிளஸ்டரிலிருந்து மெட்ரிக்குகளை சேகரிப்பது மூன்று components கொண்டது:
- Sources: இந்த guide-இல் பட்டியலிடப்பட்டவை போன்ற மெட்ரிக்குகள் வரும் இடங்கள்.
- Agents: EKS environment-இல் இயங்கும் Applications, பொதுவாக agent என அழைக்கப்படும், monitoring data-ஐ சேகரித்து இரண்டாவது component-க்கு push செய்கிறது. இந்த component-இன் எடுத்துக்காட்டுகள் AWS Distro for OpenTelemetry (ADOT) மற்றும் CloudWatch Agent
- Destinations: Monitoring data storage மற்றும் analysis solution, இந்த component பொதுவாக time series formatted data-க்கு optimized data service ஆகும். இந்த component-இன் எடுத்துக்காட்டுகள் Amazon Managed Service for Prometheus மற்றும் AWS Cloudwatch.
குறிப்பு: இந்த section-இல், configuration examples AWS Observability Accelerator-இன் relevant sections-க்கான links ஆகும். EKS metrics collection implementations-க்கான up to date guidance மற்றும் examples வழங்குவதற்காக இது செய்யப்பட்டுள்ளது.
Managed Open Source Solution
AWS Distro for OpenTelemetry (ADOT) என்பது OpenTelemetry project-இன் supported version ஆகும், users correlated metrics மற்றும் traces-ஐ Amazon Managed Service for Prometheus மற்றும் AWS Cloudwatch போன்ற monitoring data collection solutions-க்கு அனுப்ப உதவுகிறது. ADOT-ஐ EKS Managed Add-ons மூலம் EKS cluster-இல் install செய்து, இந்த page-இல் listed மெட்ரிக்குகள் மற்றும் workload traces சேகரிக்க configure செய்யலாம். AWS ADOT add-on Amazon EKS-உடன் compatible என validate செய்துள்ளது, latest bug fixes மற்றும் security patches-உடன் regularly update செய்யப்படுகிறது. ADOT best practices மற்றும் மேலும் தகவல்.
ADOT + AMP
AWS Distro for OpenTelemetry (ADOT), Amazon Managed Service for Prometheus (AMP), மற்றும் Amazon Managed Service for Grafana (AMG)-உடன் விரைவாக தொடங்க, AWS Observability Accelerator-இன் infrastructure monitoring example-ஐ பயன்படுத்தவும். Accelerator examples out of the box metrics collection, alerting rules மற்றும் Grafana dashboards-உடன் tools மற்றும் services-ஐ உங்கள் environment-இல் deploy செய்கின்றன.
EKS Managed Add-on for ADOT-இன் installation, configuration மற்றும் operation பற்றிய கூடுதல் தகவலுக்கு AWS documentation-ஐ பார்க்கவும்.
Sources
EKS மெட்ரிக்குகள் overall solution-இன் வெவ்வேறு layers-இல் பல locations-இலிருந ்து உருவாக்கப்படுகின்றன. Essential metrics section-இல் சுட்டிக்காட்டப்படும் metrics sources-ஐ சுருக்கமாக காட்டும் table இது.
Agent : AWS Distro for OpenTelemetry
AWS EKS ADOT managed addon மூலம் உங்கள் EKS cluster-இல் ADOT-ஐ install, configure மற்றும் operate செய்வதை AWS பரிந்துரைக்கிறது. இந்த addon ADOT operator/collector custom resource model-ஐ பயன்படுத்தி உங்கள் cluster-இல் multiple ADOT collectors-ஐ deploy, configure மற்றும் manage செய்ய உதவுகிறது. இந்த addon-இன் installation, advanced configuration மற்றும் operations பற்றிய விரிவான தகவலுக்கு இந்த documentation-ஐ பார்க்கவும்.
குறிப்பு: AWS EKS ADOT managed addon web console ADOT addon-இன் advanced configuration-க்கு பயன்படுத்தலாம்.
ADOT collector configuration-இல் இரண்டு components உள்ளன.
- Collector deployment mode (deployment, daemonset, etc.) உள்ளடக்கிய collector configuration.
- Metrics collection-க்கு தேவையான receivers, processors மற்றும் exporters உள்ளடக்கிய OpenTelemetry Pipeline configuration. Configuration snippet எட ுத்துக்காட்டு:
config: |
extensions:
sigv4auth:
region: <YOUR_AWS_REGION>
service: "aps"
receivers:
#
# Scrape configuration for the Prometheus Receiver
# This is the same configuration used when Prometheus is installed using the community Helm chart
#
prometheus:
config:
global:
scrape_interval: 60s
scrape_timeout: 10s
scrape_configs:
- job_name: kubernetes-apiservers
bearer_token_file: /var/run/secrets/kubernetes.io/serviceaccount/token
kubernetes_sd_configs:
- role: endpoints
relabel_configs:
- action: keep
regex: default;kubernetes;https
source_labels:
- __meta_kubernetes_namespace
- __meta_kubernetes_service_name
- __meta_kubernetes_endpoint_port_name
scheme: https
tls_config:
ca_file: /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
insecure_skip_verify: true
...
...
exporters:
prometheusremotewrite:
endpoint: <YOUR AMP WRITE ENDPOINT URL>
auth:
authenticator: sigv4auth
logging:
loglevel: warn
extensions:
sigv4auth:
region: <YOUR_AWS_REGION>
service: aps
health_check:
pprof:
endpoint: :1888
zpages:
endpoint: :55679
processors:
batch/metrics:
timeout: 30s
send_batch_size: 500
service:
extensions: [pprof, zpages, health_check, sigv4auth]
pipelines:
metrics:
receivers: [prometheus]
processors: [batch/metrics]
exporters: [logging, prometheusremotewrite]
முழுமையான best practices collector configuration, ADOT pipeline configuration மற்றும் Prometheus scrape configuration Observability Accelerator-இல் Helm Chart ஆக காணலாம்.
Destination: Amazon Managed Service for Prometheus
ADOT collector pipeline Prometheus Remote Write capabilities பயன்படுத்தி AMP instance-க்கு மெட்ரிக்குகளை export செய்கிறது. Configuration snippet எடுத்துக்காட்டு, AMP WRITE ENDPOINT URL-ஐ கவனிக்கவும்:
exporters:
prometheusremotewrite:
endpoint: <YOUR AMP WRITE ENDPOINT URL>
auth:
authenticator: sigv4auth
logging:
loglevel: warn
முழுமையான best practices collector configuration, ADOT pipeline configuration மற்றும் Prometheus scrape configuration Observability Accelerator-இல் Helm Chart ஆக காணலாம்.
AMP configuration மற்றும் usage-க்கான best practices இங்கே உள்ளது.
தொடர்புடைய மெட்ரிக்குகள் எவை?
மெட்ரிக்குகள் குறைவாக இருந்த நாட்கள் போய்விட்டன, இன்று நூற்றுக்கணக்கான மெட்ரிக்குகள் கிடைக்கின்றன. Observability first mindset-உடன் system கட்டமைக்க relevant metrics-ஐ determine செய்வது முக்கியம்.
இந்த guide உங்களுக்கு கிடைக்கும் மெட்ரிக்குகளின் வெவ்வேறு groupings-ஐ விவரிக்கிறது மற்றும் infrastructure மற்றும் applications-இல் observability கட்டமைக்கும்போது எவற்றில் கவனம் செலுத்த வேண்டும் என்பதை விளக்குகிறது. கீழே உள்ள மெட்ரிக்குகள் பட்டியல் best practices அடிப்படையில் monitoring செய்ய நாங்கள் பரிந்துரைக்கும் மெட்ரிக்குகளின் பட்டியலாகும்.
பின்வரும் sections-இல் listed மெட்ரிக்குகள் AWS Observability Accelerator Grafana Dashboards மற்றும் Kube Prometheus Stack Dashboards-இல் highlighted மெட்ரிக்குகளுக்கு additional ஆகும்.