Spring Boot Actuator Metrics with Prometheus and Grafana
Published Updated Spring Boot 12 min read
Wire Micrometer to a Prometheus scrape, get percentile latency out of http_server_requests, and build a dashboard that answers a question instead of filling a screen with gauges.
Spring Boot Actuator exposes metrics; Prometheus stores them over time; Grafana draws them. The three-way setup is short, and almost every guide to it stops at the point where a dashboard appears with numbers on it. The interesting part starts just after: which metrics are worth a panel, why your latency percentiles are wrong unless you ask for them explicitly, and how a single careless tag can multiply your time series until Prometheus starts dropping data.
Written against Spring Boot 3.2, Micrometer 1.12, Prometheus 2.x and Grafana 10.
What Actuator actually contributes
Actuator does not implement metrics. It ships Micrometer, a facade over a metrics backend in the same way SLF4J is a facade over a logging backend. Your code, and every instrumented Spring component, records against Micrometer’s API, and a registry on the classpath decides where the numbers go. Add the Prometheus registry and Micrometer starts keeping data in the shape Prometheus wants and publishes it on an HTTP endpoint for scraping.
Nothing pushes. Prometheus pulls, on an interval, over HTTP. That single fact explains most of the setup problems further down, if the Prometheus process cannot reach your application’s endpoint, there is no data, and the application has no idea anything is wrong.
Two dependencies
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-actuator</artifactId>
</dependency>
<dependency>
<groupId>io.micrometer</groupId>
<artifactId>micrometer-registry-prometheus</artifactId>
<scope>runtime</scope>
</dependency>
The registry is runtime scope on purpose. You never import from it. Its presence on the
classpath is the whole configuration. Micrometer’s auto-configuration finds it and enables the
matching endpoint.
Exposing the endpoint
Actuator endpoints are enabled by default but only health is exposed over HTTP. Those are two
separate settings and conflating them is the first thing that goes wrong:
management.endpoints.web.exposure.include=health,info,prometheus
management.endpoint.health.show-details=when-authorized
Restart and the scrape target exists:
$ curl -s localhost:8080/actuator/prometheus | head -20
# HELP jvm_memory_used_bytes The amount of used memory
# TYPE jvm_memory_used_bytes gauge
jvm_memory_used_bytes{area="heap",id="G1 Eden Space"} 2.4117248E7
jvm_memory_used_bytes{area="heap",id="G1 Old Gen"} 1.1534336E7
# HELP http_server_requests_seconds Duration of HTTP server request handling
# TYPE http_server_requests_seconds summary
http_server_requests_seconds_count{method="GET",outcome="SUCCESS",status="200",uri="/api/notes"} 3.0
http_server_requests_seconds_sum{method="GET",outcome="SUCCESS",status="200",uri="/api/notes"} 0.0412
This is a text format, not JSON, and it is a snapshot: every scrape returns current values, and the history lives entirely in Prometheus. That is why hitting the endpoint by hand tells you the instrumentation works but nothing about a trend.
Note the uri label. It is the templated path, /api/notes/{id}, never /api/notes/42. Spring
tags the route, not the request. Hold on to that; it matters in the last section.
The metrics worth knowing by name
Out of several hundred series, four families carry most of the diagnostic value:
http_server_requests_seconds: request count and duration, tagged byuri,method,statusandoutcome. Nearly every question about “is the app healthy” is a query over this one.jvm_memory_used_bytes/jvm_gc_pause_seconds, heap by region and garbage collection pauses. Rising old-gen usage that never drops after a collection is the classic leak signature.hikaricp_connections_pendingandhikaricp_connections_acquire_seconds, threads waiting for a database connection. A latency spike with flat CPU is very often this.process_cpu_usageandsystem_cpu_usage, the process against the machine it shares.
Percentiles do not exist until you ask
By default http_server_requests_seconds is published as a summary: a count and a total. From
those two numbers you can compute a mean, and a mean latency hides exactly the problem you are
looking for. The slow tail is where users live.
Turn on histogram buckets for that meter:
management.metrics.distribution.percentiles-histogram.http.server.requests=true
management.metrics.distribution.slo.http.server.requests=50ms,100ms,250ms,500ms,1s
The endpoint now emits http_server_requests_seconds_bucket series, and Prometheus can compute a
quantile across instances:
histogram_quantile(0.95,
sum(rate(http_server_requests_seconds_bucket[5m])) by (le))
The distinction that trips people: Micrometer can also publish percentiles=0.95 directly, which
computes the quantile inside each application instance. Those pre-computed values cannot be
aggregated. Averaging two instances’ p95 is not the p95 of the pair. Histogram buckets can be
summed, which is why they are the right choice as soon as you run more than one replica.
Running Prometheus and Grafana
Prometheus needs to know where to scrape:
scrape_configs:
- job_name: notes-api
metrics_path: /actuator/prometheus
scrape_interval: 10s
static_configs:
- targets: ['host.docker.internal:8080']
services:
prometheus:
image: prom/prometheus:v2.53.0
ports: ["9090:9090"]
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
grafana:
image: grafana/grafana:10.4.3
ports: ["3000:3000"]
environment:
GF_SECURITY_ADMIN_PASSWORD: admin
host.docker.internal is how a container reaches a process on the host under Docker Desktop. On
plain Linux that name does not resolve: either run the compose file with
extra_hosts: ["host.docker.internal:host-gateway"], or put the application in the same compose
project and use its service name. Getting this wrong is the single most common failure, and its
symptom is a target sitting at DOWN in localhost:9090/targets with a connection-refused
error. Check that page before touching Grafana.
In Grafana, add Prometheus as a data source at http://prometheus:9090, the service name, not
localhost, because Grafana is also in a container. Then import dashboard 4701 (“JVM
Micrometer”) for a serviceable JVM view in about fifteen seconds.
Three queries worth a panel
Imported dashboards are a starting point, not a monitoring strategy. These three answer questions someone actually asks during an incident:
# request rate by route
sum(rate(http_server_requests_seconds_count[5m])) by (uri)
# error ratio, 0 to 1
sum(rate(http_server_requests_seconds_count{outcome="SERVER_ERROR"}[5m]))
/ sum(rate(http_server_requests_seconds_count[5m]))
# p95 latency for one route
histogram_quantile(0.95,
sum(rate(http_server_requests_seconds_bucket{uri="/api/notes"}[5m])) by (le))
rate() over a counter, never the counter itself: the raw number only ever climbs and resets on
restart, so graphing it tells you the process has been up, not that it is serving.
Your own metrics
Inject MeterRegistry and register instruments in the constructor rather than per call. Creating
a meter is a map lookup, but the tags are evaluated every time:
@Service
public class NoteService {
private final Counter created;
private final Timer searchTimer;
public NoteService(MeterRegistry registry) {
this.created = Counter.builder("notes.created")
.description("Notes successfully persisted")
.register(registry);
this.searchTimer = Timer.builder("notes.search")
.publishPercentileHistogram()
.register(registry);
}
public Note create(Note note) {
Note saved = repository.save(note);
created.increment();
return saved;
}
public List<Note> search(String q) {
return searchTimer.record(() -> repository.search(q));
}
}
Micrometer names meters with dots and Prometheus renames them on the way out —
notes.created becomes notes_created_total. Search for the Prometheus spelling when a metric
seems to be missing.
The mistake that takes a monitoring stack down
Every distinct combination of label values is a separate time series. Tag a metric with something unbounded (a note id, a user id, a raw URL, an exception message) and you create a series per value. Prometheus holds active series in memory; a few thousand is unremarkable, a few million is an outage, and the thing that falls over is your monitoring rather than your application.
This is why Spring tags uri with the route template. If you build your own tag from a path
variable you have undone that protection:
// wrong: one time series per id, forever
Counter.builder("notes.fetched").tag("id", id.toString()).register(registry);
// right: bounded set of values
Counter.builder("notes.fetched").tag("result", found ? "hit" : "miss").register(registry);
Ask of every tag: how many distinct values can this have? If the answer is “as many as we have rows”. It is not a tag. It is a log line.
Frequently asked questions
Why is /actuator/prometheus returning 404?
The endpoint is enabled but not exposed. Add
prometheus to management.endpoints.web.exposure.include: health is the only endpoint exposed
over HTTP by default.
Do I need Spring Boot Actuator if I already have Micrometer?
Yes, in practice. Actuator provides the auto-configuration that creates the registry, binds the JVM and web instrumentation, and publishes the scrape endpoint. Micrometer alone gives you the API and nothing wired up.
Why is my Prometheus target DOWN?
Almost always a networking boundary. A Prometheus container
cannot reach the host’s localhost; use host.docker.internal with a host-gateway entry on
Linux, or put both in one compose network. Read the error on /targets rather than guessing.
Where do latency percentiles come from?
Not from the defaults. Set
management.metrics.distribution.percentiles-histogram.http.server.requests=true to publish
histogram buckets, then compute quantiles with histogram_quantile() in PromQL.
Should I use percentiles-histogram or percentiles?
Histogram, once you run more than one instance. Client-side percentiles are computed per instance and cannot be aggregated across them; buckets can be summed.
My custom metric does not appear in Prometheus.
Two likely causes: it has not been recorded yet
(a meter with no observations is not published), or you are searching for the Micrometer name.
notes.created is exported as notes_created_total.
How do I secure the endpoints?
Put them on a separate port with
management.server.port=8081 and leave that port off the public load balancer, or protect
EndpointRequest.toAnyEndpoint() with Spring Security. Exposing /actuator/env or
/actuator/heapdump publicly is a data leak, not an inconvenience.
Does scraping cost anything?
Serialising a few thousand series takes single-digit milliseconds and allocates. A 10-second interval is fine; a 1-second interval across many instances is measurable and rarely tells you anything the 10-second one did not.
Can Prometheus keep months of history?
Local retention defaults to 15 days and it is not built to be a long-term store. For longer horizons, remote-write to something designed for it (Thanos, Mimir or Cortex) rather than growing the local volume.
What about /actuator/metrics?
A JSON endpoint for reading one meter by hand, useful for a quick check. It is not the scrape endpoint and Prometheus does not read it.
Where should I go next?
The Actuator endpoint reference covers the
other endpoints (health groups, /loggers, auditing) and the
Spring Boot REST API guide is the
application these metrics are measuring.