Showing posts with label Monitoring. Show all posts
Showing posts with label Monitoring. Show all posts

Monday, May 12, 2025

Prometheus - Write a Query through PromQL - Part 2

May 12, 2025 0

 

PromQL:
PromQL expressions has classified as follows:
Instant Vector
Range Vector
Scalar
String
Instant and Range vector has a basic syntax consist of metric name along with label matchers. A metric name in Prometheus is a label named __name__ under the value. 
example:
{__name__="prometheus_build_info"}
PromQL will query directly by name [prometheus_build_info]

Label matchers are defined using a label key, a label matching operator, and the label value. Some possible label-matching operators:
=: The label value matches the specified string
!=: The label value does not match the specified string
=~: The label value matches the regex in the specified string
!~: The label value does not match the regex in the specified string

Instant Vector are selectors that select the data at this instant using a recent available data for the time series which you are selected. 
Range Vector is retriving data over a given time range. This time range is specified as duration followed by unit.
The following are valid units for a duration:

ms: Milliseconds
s: Seconds
m: Minutes
h: Hours
d: Days (assumes a day is always 24 hours)
w: Weeks (assumes a week is always 7 days)
y: Years (assumes a year is always 365 days)

These range vector selectors are often more useful when aggregating data to perform an analysis to drive results such as the number of requests per second a service experienced over a particular time range.
Prometheus has two primary API endpoints that are used to retrive time series data /api/v1/query and /api/v1/query_range.
Note : Both instant vectors and range vectors are valid for the /query endpoint, but only instant vectors are valid for the /query_range endpoint.
Offsets:
It will allow us to pull the time series data from the past based on offset values.
Example:
my_cool_custom_metric offset 5m
Group Modifier and Vector Matching:
Prometheus has provided ways to join the labels of different matrics together when using the group_left and group_right query modifiers. This option allows for many-to-one and one-to-many matching of vectors.
Logical and set binary operators
PromQL used the boolean operators of and,or and unless query operators
Example:
vector1 and vector2
It requires that for every series in vector1, there must be an exactly matching label set in vector2 (besides the __name__ label). If there is no match for the label set of a series in vector1 inside of vector2, then that series is not included in the results.  

Sunday, May 4, 2025

Prometheus - Time Series Database - Part 1

May 04, 2025 0

 


Prometheus Installation:

We can install the Prometheus in different ways.
  •  Install through source
  • Docker container
  • Install through script
  • Install through package manager
Prometheus with a metrics comes from node_exporter process from remote hosts. Each metrics has associated with one or more-time series data on it.
Events are point-in-time of record of action occurring during that time.
Prometheus is using a pull-based model for extracting the data from systems it monitors whereas Graphite and Nagios are using push-based model which push the metrics to a remote system.
Components of Prometheus Stack:
  • Prometheus Alert manager for routing an alert coming from Prometheus 
  • Exporters: It has multiple exporter components depends upon a device such as Node Exporter.
  • Grafana: It used to visualize metrics from Prometheus through dashboards.
Prometheus has four main components such as time series database [TSDB], scrape manager, rule manager and Web UI.
Time Series Database:
   It is a special kind of database that is optimized for storing data points in a time series manner. The core part of TSDB are head blocks [Write-ahead log] and data format (block, chunks, and indices). The head block is an entry point for samples being scraped and stored in TSDB. 
The chunk gets appending into head block, until chunk hits it sample limit until stays in memory. It will add into WAL at first before added into memory, so we will preserve the data if server is in panic or reboot at automatically.
To increase resource efficiency, some wizardry is done to the chunk data so that it only stores the direct value of the first timestamp (t0) and first value (v0). Subsequent timestamps are set to the delta of the prior timestamp, so starting at t2, all timestamps are the delta of a delta. Similarly, subsequent values are compared to their prior value using a bitwise XOR operator. This just means that the difference between the samples is stored, and if they’re the same, then 0 is stored.
Scrape Manager:
  The scrape manager is a part of Prometheus that handles pulling metrics from applications and exporters by performing scraping and maintaining an internal list of what things should be scraped by Prometheus.
Rule Manager:
  The rule manager is part of Prometheus that handles evaluating alerts and recording the rules as per Prometheus. It is handling evaluating rule groups compromised of alerts and recording rules on regular intervals. 
Web UI/API:
  The Web UI/API is the portion of Prometheus that you can access via your browser. The Prometheus UI/API is a powerful REST API. It provides integration with other graphical tools such as Grafana. 
Alert Manager:
  It is responsible to send an alert to Slack, PagerDuty and other alerting destinations.  Alert Manager handles all of that owns through routing tree-based workflow. 





Monday, March 3, 2025

Datadog Introduction

March 03, 2025 0

 


Datadog Monitoring Tool:
    Observability is essential for managing modern infrastructure and applications. It brings together real-time metrics from servers, containers, databases, and applications. It achieves this with end-to-end tracing. That’s not all that it can do. It comes up with helpful alerts and fascinating visualizations, offering full-stack observability.

Observability of three metrics [Rate, errors and duration] provide a well rounded view of service performance.
Rate : Monitoring a HTTP and API calls of your services
We will get to know the service over load for monitoring the HTTP and API calls. We will take action if any spikes or sudden drops in the rate of requests. It could indicate issues such as sudden traffic surges, DDos attacks or failures in the upstream.
Errors: Track how many of those request fail.
The error could be server error, database error or failed API calls. We can quickly identify issues with our application or backend system while tracking these errors.
Duration: Measure how long those requests take with latency
High latency will degrade a user experience particularly for real time or interactive applications. Monitoring of duration will allow us to detect a performance bottlenecks before they affect our users.

Below are few of the Datadog monitoring types and we can create and use it.

APM: Monitor application performance monitoring (APM) metrics or trace queries.
Metric: Compare values of a metric with a user-defined threshold.
Logs: Alert when a specified type of log exceeds a user-defined threshold over a given period of time.
Database Monitoring: Monitor query execution and explain plan data gathered by Datadog.
Error Tracking: Monitor issues in your applications gathered by Datadog.
Real User Monitoring: Observe user behavior and monitor frontend performance.
Synthetic Monitoring: Simulate user actions to test API endpoints or website functionality.
Anomaly: Detect anomalous behavior for a metric based on historical data.
Cloud Network Monitoring: Monitor cloud-specific network configurations and traffic.

Infrastructure Monitoring: 
Datadog can monitor the performance and health check of our entire infrastructure. This includes servers, containers, databases, and cloud services. It provides:
  • Metrics collection
  • Visualizations and dashboards
  • Alerting
  • Anomaly [behaves differently than usual] monitoring
  • Infrastructure maps
  • Logs and traces integration
  • Automation
Application Performance Monitoring (APM)
Datadog offers APM functionality to monitor and optimize the performance of your applications. It provides detailed visibility into application code, dependencies, and performance bottlenecks. With it, you can track response times, error rates, and throughput. What’s more, you can gain visibility into the performance of individual requests.

Distributed Tracing
Distributed tracing, Datadog allows teams to trace requests as they flow through your complex, distributed systems. It helps you to:
  • Identify latency issues
  • Understand dependencies between services
  • Troubleshoot performance problems across microservices architectures
  • You can see how you can easily identify the root causes of application performance issues. It collects data moving between services.
Log Management
Datadog enables centralized log management. This allows you to collect, index, search, and analyses logs from various sources. You can aggregate Datadog logs from multiple systems and applications. You can set up alerts based on log events. Also, you can collect the customized logs from the servers.
Real-time Metrics and Dashboards
 Datadog provides real-time metrics and customizable dashboards. A Datadog dashboard helps to visualize and monitor the health and performance of our systems. You can create visualizations, charts, and graphs to:
  • Track key metrics
  • Set up alerts based on thresholds
  • Share dashboards with your team
Collaboration and Notifications
Datadog offers collaboration features that allows teams to work together effectively. You can annotate and share graphs, dashboards, and alerts. You can even set up notifications via email, SMS, or third-party integrations. Want to integrate incident management tools like Slack, Jira, and PagerDuty? No problem. You can even collaborate on troubleshooting and resolving issues.
Integration and Extensibility
Datadog integrates with a wide range of tools and services. As you can imagine, this makes it easy to collect data from various sources. You can integrate Datadog with all the popular cloud platforms. Don’t stop there, you can integrate it with all your existing workflows.

Tuesday, October 8, 2024

Prometheus Introduction

October 08, 2024 0

 



Prometheus:

Prometheus is an open-source toolkit written in Go that is designed to be a fully featured solution for monitoring and alerting.

Prometheus is responsible for collecting metrics data and storing it in an efficient time series database, while user is able to query that time-series data and configure alerting for real-time updates.

It actually will use the built-in database within it. We can transfer the data into external storage or keep within it.

It uses a multidimensional data model that accommodates time-series data. This data is associated with the timestamp and options key-value pairs.

Prometheus uses a HTTP pull model to retrieve stored time-series data.

Prometheus Components:

Prometheus server is a responsible for collecting a data from exporter or scraping a data from targeting machines and storing this data into time-series database.

Various client libraries are supported for programming languages including Rust, Python, Java, Ruby and Go. These libraries aid in instrumenting application mode.  Special exporters are also supported for exposing metrics from systems that cannot directly use Prometheus metrics such as Graphite, StatsD and other third parity software.

Alertmanager is used to handling alert within Prometheus.

Service discovery mechanisms such as Kubernetes native service discovery, DNS and file_sd are supported and discover  and begin monitoring new targets automatically.

The Push gateway is a separate component which used to collect metrics from short live which cannot be capurate by usual monitoring methods.

PromQL is an addition component used to provide a built-in expressive query language for querying and aggregating time-series data within Prometheus.

Prometheus pull a data from specfic targets through collecting metrics from http end points.  The main configuration file is prometheus.yml.  We can setup Prometheus to collect metrics on itself or pull a data and monitor own health.

PromQL expression are entered in the promethus expression browser [http://localhost:9090/graph]. We can enter any expression in order to render the result in a table or graph the result over time.

Exports are tools that help export metrics from third-parity system into Prometheus for consumption and aggregation.

Alert manager is efficient tool for collecting a alerts and manage it. It can able to redirect the notification to slack, PagerDuty, OpsGenie, Telegram, Microsoft Teams, WebEx, Discord and other receivers.