Redis Monitoring: The Metrics That Matter, and What Your Client Isn't Telling You

Last updated
August 20, 2026

Most Redis monitoring advice is a list of fields from INFO and a threshold to alert on. That is a reasonable start, and it is where nearly every guide stops. Three things get lost. Redis's own documentation says the memory metric everyone alerts on is not the one that measures fragmentation. The subsystem behind its most useful latency command ships switched off. And the complaint that brings most people here — Redis says it is healthy and my application says it is slow — is invisible to all of it.

This article covers what to monitor on the server, which commands actually answer which question, how the tools compare, what managed providers take away, and the layer that server-side monitoring structurally cannot see. Server behaviour below is checked against the Redis and Valkey documentation and, where the two disagree, against their source. Vendor pricing and product details reflect published terms as of August 2026 and change often.

What to Monitor in Redis

Redis exposes its internal state through a single command, INFO, which returns a few hundred fields grouped into sections. You do not need most of those fields. The Redis metrics worth alerting on number around twenty, and they fall into seven groups: memory, keyspace effectiveness, throughput and latency, connections, CPU, replication, and durability.

Field Section What it tells you

used_memory

memory

Bytes allocated by Redis through its own allocator. Compare against maxmemory, not against the machine's RAM

used_memory_rss

memory

Resident set size as the operating system sees it — the number top reports

allocator_frag_ratio

memory

The real fragmentation signal. See the next section — it is not mem_fragmentation_ratio

maxmemory_policy

memory

Which eviction policy is active. noeviction means writes start failing instead of data being dropped

evicted_keys

stats

Keys dropped because of the maxmemory limit. For a cache this is normal; for a store it is data loss

expired_keys

stats

Keys removed by TTL. Rising here is the system working as designed

keyspace_hits / keyspace_misses

stats

Successful and failed lookups in the main dictionary — the inputs to hit ratio

instantaneous_ops_per_sec

stats

Current throughput. Useful mostly as a shape, to correlate against latency

latency_percentiles_usec_<cmd>

latencystats

Per-command p50, p99 and p99.9 in microseconds. On by default, but not in a bare INFO — see below

latest_fork_usec

stats

Duration of the last fork, in microseconds. Long forks stall the server and are a classic hidden latency source

connected_clients

clients

Client connections, excluding replicas

blocked_clients

clients

Clients waiting on a blocking call such as BLPOP or BZPOPMIN

rejected_connections

stats

Connections refused at the maxclients limit. Any non-zero value is worth an alert

used_cpu_user_main_thread / used_cpu_sys_main_thread

cpu

Command execution is single-threaded, so the main thread saturating is the real ceiling — not whole-process CPU. Watch both: which dominates depends on the workload, and small commands at high rates push system time up. Linux only

master_link_status

replication

On a replica, up or down. Pair with master_last_io_seconds_ago

sync_full

stats

Full resyncs with replicas. Repeated full resyncs usually point to an undersized replication backlog, a primary restarted from scratch, or a replica restarted from AOF. Since Redis 4.0 a failover should not force one

rdb_last_bgsave_status / aof_last_write_status

persistence

Whether the last save and the last AOF write succeeded. Silent failure here is how backups turn out to be missing

Two pairs in that table are routinely conflated. evicted_keys and expired_keys are different events: expiry is a TTL doing its job, eviction is Redis at its memory ceiling discarding data to accept a write. And there are two AOF status fields — aof_last_bgrewrite_status for the last rewrite and aof_last_write_status for the last write. They fail independently, so alert on both. For how the durability side works, see Redis persistence; for failover behaviour, Redis Sentinel.

Redis Memory Metrics: Why mem_fragmentation_ratio Misleads

The standard advice is to alert when mem_fragmentation_ratio goes above 1.5. Redis's own INFO documentation says, of allocator_frag_ratio, that it "is the true (external) fragmentation metric (not mem_fragmentation_ratio)." The parenthesis is in the original.

The reason is that mem_fragmentation_ratio is simply used_memory_rss divided by used_memory, and the documentation is explicit that this "doesn't only includes fragmentation, but also other process overheads... and also overheads like code, shared libraries, stack, etc." (the grammatical slip is in the original). A process that has loaded a large shared library has a raised ratio and no fragmentation problem at all.

The thresholds Redis itself uses are visible in MEMORY DOCTOR, and they are conjunctions rather than single numbers. It reports high RSS only when mem_fragmentation_ratio exceeds 1.4 and the fragmentation delta exceeds 10 MB; it reports genuine allocator fragmentation when allocator_frag_ratio exceeds 1.1 and that delta exceeds 10 MB. Below 5 MB of total allocation it declines to diagnose at all, replying only that the instance is using too little memory for its issue detector to be used. The documentation makes the same point in prose: "when the total fragmentation bytes is low (few megabytes), a high ratio (e.g. 1.5 and above) is not an indication of an issue."

A ratio below 1 is the case worth understanding rather than alerting on. It means resident memory is smaller than what Redis believes it has allocated, which follows from the definition if some pages are not resident — the usual explanation being that the operating system has swapped them out, which is catastrophic for a system whose entire premise is RAM speed. This inference is near-universal in monitoring guides but neither redis.io nor valkey.io actually documents it, and what they do document is a better test anyway: check /proc/<pid>/smaps for Swap: entries, or watch the si and so columns of vmstat 1. Those measure swapping directly instead of inferring it — though both assume a shell on the host, which on managed Redis you will not have.

Eviction Policies: Eight, or Ten on Redis 8.6

When used_memory reaches maxmemory, what happens next is entirely determined by maxmemory-policy. The familiar list has eight entries. It now has ten: Redis 8.6 added volatile-lrm and allkeys-lrm, where LRM is Least Recently Modified — the timestamp updates on writes but not on reads, which is what you want when the interesting question is what has gone stale rather than what is unpopular. Valkey does not have LRM as of this writing, so if you run both engines this is a real divergence. The Redis eviction policy page covers the full set.

Two behaviours in that mechanism surprise people. The default policy is noeviction, under which Redis does not discard anything and instead returns an error on commands that would allocate. The configuration file gives examples that apply under any policy once no suitable key can be evicted: SET, INCR, HSET, LPUSH, SUNIONSTORE, SORT with a STORE argument, and EXEC where the transaction includes any command that requires memory. Reads keep working, so the failure presents as a write outage rather than a slowdown. And every volatile-* policy behaves exactly like noeviction when no key carries a TTL, which is how a cache configured for volatile-lru can start refusing writes while the operator is looking at a dashboard showing eviction is enabled.

Hit ratio is arithmetic on two of the counters above: keyspace_hits / (keyspace_hits + keyspace_misses). Redis's OSS documentation does not state that formula outright, though Redis Software's observability guide uses the structurally identical expression and alerts below 50%. Treat it as the standard derivation it is, and read it alongside evicted_keys — a falling hit ratio with rising evictions is a sizing problem, whereas a falling hit ratio with flat evictions is usually a cache miss pattern problem in the application.

INFO, MONITOR, SLOWLOG and LATENCY

Four commands answer four different questions, and using the wrong one is how monitoring turns into an outage.

INFO

A bare INFO does not return everything. Redis defines a default section set, and commandstats and latencystats are not in it — you have to ask for them by name, or use INFO all. This trips up a lot of monitoring setups that parse a bare INFO and then wonder why per-command data never appears. Note that errorstats is in the default set.

commandstats is where per-command cost lives, as cmdstat_get:calls=...,usec=...,usec_per_call=.... Read usec carefully: it is documented as CPU time consumed, not wall-clock latency, so it will not show you time spent waiting.

MONITOR

MONITOR streams back every command the server processes. It is superb for answering "what is my application actually sending" and it must not be left running. Redis publishes its own benchmark for the cost, hedged in the documentation as "totally unscientific". In that benchmark SET ran at 95,419.85 requests per second normally and 41,823.50 with a single MONITOR client attached, and the documentation concludes that "running a single MONITOR client can reduce the throughput by more than 50%. Running more MONITOR clients will reduce throughput even more."

Throughput is not even the worst of it. Monitors are ordinary clients, and Redis ships client-output-buffer-limit normal 0 0 0 — no limit at all. Attach a monitor over a slow link to a busy instance and the server's output buffer for that client grows unbounded, which is a memory-exhaustion path rather than a slowdown. Bound it with timeout, redirect to a file rather than a terminal, and inside an interactive session leave monitor mode with RESET.

Two exclusions matter if you are relying on it for an audit trail rather than debugging. Administrative commands are not logged at all, and QUIT is not logged. AUTH is a special case: since 6.2.4 it appears in the stream, but with its arguments redacted. It is not invisible, and it is not readable.

SLOWLOG

The slow log records commands whose execution exceeded slowlog-log-slower-than, which defaults to 10,000 microseconds — 10 ms, not 10 seconds, a units error worth double-checking in your own config. It keeps slowlog-max-len entries, default 128, in memory. Read it with SLOWLOG GET, which returns the latest ten entries unless you pass a count, or -1 for all of them.

And then the caveat that matters more than anything else in this section. From the configuration file: the execution time "does not include the I/O operations like talking with the client, sending the reply and so forth, but just the time needed to actually execute the command."

Read that precisely, because it is narrower than it is usually reported. Building the reply happens inside the command and is counted — an HGETALL walks the hash, and a GET of a large value copies it, both on the main thread, both measured. What falls outside the measurement is the socket write. The gap this opens is one of cost per byte rather than of category: assembling a reply in memory is perhaps two orders of magnitude cheaper per byte than pushing it through a socket, so a command can build several megabytes in a few hundred microseconds, stay comfortably under the 10 ms threshold, and then take milliseconds to reach the client. Nothing in the slow log records that second part. The HGETALL versus HSCAN trade-off is the usual culprit, alongside unbounded key scans.

LATENCY

There are two separate latency subsystems in Redis and they are constantly confused with each other.

The latency monitor records spikes against named events — command, fork, expire-cycle, eviction-cycle, aof-fsync-always and others — in 160-element time series. It is disabled by default: latency-monitor-threshold ships at 0, because, as the configuration file explains, "collecting data has a performance impact, that while very small, can be measured under big load." Turn it on with CONFIG SET latency-monitor-threshold 100, then read it with LATENCY LATEST, LATENCY HISTORY, LATENCY GRAPH and LATENCY DOCTOR. Clear it with LATENCY RESET.

Extended latency tracking is a different thing entirely and it is on by default. latency-tracking defaults to yes on Redis and Valkey OSS — though not everywhere, ElastiCache ships it as no — and populates the latencystats section with per-command percentiles, by default p50, p99 and p99.9, configurable through latency-tracking-info-percentiles. LATENCY HISTOGRAM returns the full cumulative distribution.

The asymmetry that catches people: LATENCY RESET clears the latency monitor's time series and does nothing to the histograms. Those are cleared by CONFIG RESETSTAT. Two subsystems, two configuration flags, two resets. Because these are percentiles rather than averages, latency versus throughput is worth reading on why the average is the least useful number here, and t-digest on computing percentiles across shards without averaging them.

From the shell, redis-cli --latency pings 100 times a second and reports min, max and average in milliseconds; --latency-history restarts the sampling window every 15 seconds so you can see drift. It is worth being clear about which redis-cli flag measures which part of a request before you trust any of them, because --intrinsic-latency is the one people misuse. It does not connect to Redis at all — run on the server itself, it measures the largest interval the kernel withholds CPU from a userspace process. The documentation warns that the test "is CPU intensive and will likely saturate a single core", which is worth weighing before running it on a host that is already struggling. Its virtualised example measures 9.7 ms and concludes: "this means that we can't ask better than that to Redis." That number is the floor. Below it, tuning Redis is pointless and the problem is the host.

# four questions, four commands. Add -h/-p/-a/--tls for a remote or managed server.
redis-cli INFO all | grep -E 'used_memory:|^maxmemory:|maxmemory_policy|allocator_frag_ratio|evicted_keys|keyspace_(hits|misses)'
redis-cli SLOWLOG GET 10                     # which commands execute slowly?
redis-cli LATENCY LATEST                     # which events spiked? (needs the monitor enabled)
timeout 5 redis-cli MONITOR > /tmp/mon.log   # what is the app sending? Always bound it.

# two shell modes that are not the LATENCY subsystem
redis-cli --latency-history -i 5             # is latency drifting?
redis-cli --intrinsic-latency 30             # run ON the host; saturates one core for 30s

Two of those assume a shell on the Redis host, which on a managed service you will not have — and on ElastiCache Serverless three of the four are blocked outright. The managed section below sets out what each provider takes away.

Valkey 8.1: COMMANDLOG Replaces the Slow Log

If you run Valkey rather than Redis, the slow log has been superseded. Valkey 8.1 introduced COMMANDLOG, and its configuration file states plainly that slowlog-log-slower-than and slowlog-max-len "are still supported but deprecated" in favour of the new parameters.

The reason this matters is not tidiness. COMMANDLOG records three kinds of event, not one:

  • SLOW — execution time in microseconds. The old slow log, unchanged.

  • LARGE-REQUEST — request size in bytes, for "commands that send excessive data to the server."

  • LARGE-REPLY — reply size in bytes, with the documentation naming KEYS and HGETALL as the examples.

LARGE-REPLY is the genuinely new capability, and it is precisely the blind spot described in the previous section: the reply that costs little to build and a great deal to ship. Note that Valkey lists KEYS and HGETALL alongside a GET of a substantial value, so this is not only about commands that are cheap to execute — it is about reply size, which the execution-time log does not measure at all. Configuration is commandlog-execution-slower-than, commandlog-request-larger-than and commandlog-reply-larger-than, the latter two defaulting to 1 MiB. Valkey warns that enabling reply tracking has measurable overhead when I/O threads are in use, so it can be switched off with -1.

Redis Monitoring Tools Compared

Nearly all of them are, at bottom, a scheduler around INFO — including the provider-native services, which AWS and Microsoft both document as deriving their metrics from it. Several also issue SLOWLOG, LATENCY or CONFIG GET on each scrape, and the providers add a handful of host-level figures Redis cannot report about itself, such as CPU and network utilisation of the underlying instance. What mostly differs is how each collects, what it costs, and what it does with the data.

Tool Collection Cost model Best for

Redis Insight

Direct connection, desktop or browser

Free

Interactive inspection, keyspace browsing, profiling a running instance. Not an alerting system

Prometheus + redis_exporter

Separate exporter process, scraped. Default port 9121; runs local or remote

Free, self-hosted

The default open-source answer. Pairs with Grafana dashboards and Alertmanager

Grafana

Visualization over Prometheus, or Grafana Alloy's built-in Redis component

Free OSS, paid Cloud

Dashboards on top of whatever is collecting

Datadog

Agent integration

Per host. Redis is a standard integration, so its metrics — including the optional command stats — are not billed as custom metrics

Teams already standardised on it; correlation across the whole stack

Dynatrace

OneAgent, auto-discovery

Platform subscription; full-stack billed per memory-GiB-hour, infrastructure-only per host-hour

Large estates where manual instrumentation does not scale

New Relic

Agent integration

Per GB ingested, plus per full-platform user

APM-first shops wanting Redis in the same trace view

Zabbix

Template plus agent

Free, self-hosted

Existing Zabbix estates, on-premises, air-gapped environments

CloudWatch / Azure Monitor / Cloud Monitoring

Provider-native, nothing to deploy

Platform metrics are free on both; you pay for alarms and alert rules

ElastiCache, Azure Cache and Azure Managed Redis, especially where nothing may be deployed alongside the cache

Prometheus and the Redis Exporter

Two notes on the Redis exporter, since it is the one most teams end up running. It is oliver006/redis_exporter, a community project — it is not maintained by Redis Ltd, whose own Prometheus path for Redis Software is a separate, built-in endpoint. And its metric names are not a mechanical transformation of the INFO field names. used_memory becomes redis_memory_used_bytes predictably enough, but sync_full becomes redis_replica_resyncs_full, and sync_partial_err becomes redis_replica_partial_resync_denied. Alert rules written by guessing the name will silently never fire.

Monitoring Managed Redis

On managed Redis the provider owns the host, which changes both what you can collect and how you collect it. Amazon ElastiCache publishes to CloudWatch, Azure Cache for Redis and its successor Azure Managed Redis to Azure Monitor, and Google Memorystore to Cloud Monitoring — in each case without an agent, because you cannot install one on the node itself. That constraint is narrower than it sounds: most tools in the previous section can still be pointed at a managed endpoint from a host you control, so provider-native metrics are a floor rather than a ceiling. Dynatrace is the notable exception: its own extension documentation requires that Redis be listening on the OneAgent host's loopback interface, so for ElastiCache it reads CloudWatch instead. An exporter aimed at a managed endpoint also loses every CONFIG-derived metric, for the reason in the next paragraph.

A currency note while you are choosing dashboards: Microsoft has published a retirement timeline for the Azure Cache for Redis SKUs and now directs both new and existing workloads to Azure Managed Redis, so check which product you will still be running before you build against the older one.

What you lose is more interesting than what you gain. MEMORY DOCTOR is documented as unavailable on Redis Software and Redis Cloud, so the threshold logic described earlier has to be reimplemented from raw fields. CONFIG is unavailable on ElastiCache and disabled on Azure Cache for Redis alike, which means CONFIG SET latency-monitor-threshold is simply not available to you — and since the latency monitor is off by default, that leaves the subsystem permanently unreachable on those platforms. MONITOR is restricted less uniformly: ElastiCache blocks it on serverless caches but allows it on node-based clusters. Serverless is the harsher tier generally — it also removes SLOWLOG, every LATENCY subcommand, COMMANDLOG and the whole MEMORY family, which, together with MONITOR, takes away three of the four commands in the previous section. Provider metric names are also their own vocabulary rather than INFO field names, so dashboards do not port between a self-hosted instance and a managed one.

The Java connection details differ per provider too: AWS ElastiCache, Azure Cache and Google Cloud Memorystore each have their own. In cluster mode, remember that per-node metrics need to be read per node; a cluster-wide average will hide the one hot shard that is actually causing your incident.

What Server-Side Redis Monitoring Cannot See

Everything above measures the server. None of it measures your application's experience of the server, and those are not the same thing — the same split that runs through Valkey and Redis best practices for Java. Four failure modes are effectively invisible from the Redis side:

  • Connection pool exhaustion. The pool is in your JVM. Redis sees a healthy, modest number of connections; your threads are queuing for one. connected_clients looks fine because it is fine.

  • Latency as the application measures it. The slow log excludes the socket write by design, and knows nothing of the network. Round trip and reply transfer are outside it, and all inside your latency budget.

  • Reconnect and retry storms. A brief network partition or a failover produces a burst of reconnections and retried commands. From the server this is a small blip in connected_clients; from the client it is seconds of stalled work.

  • Local cache effectiveness. If you run a near cache or client-side caching, its hits never reach Redis at all. They are invisible to keyspace_hits by definition — the better it works, the less the server sees.

Closing that gap means instrumenting the client, and the raw signals are not exclusive to any one library: Lettuce can record command latency through Micrometer, and Jedis exposes pool counters through commons-pool2 and client-side-cache statistics through CacheStats and Cache.getSize(). What none of them ships is those signals assembled into one per-node metric set and wired to an external system without you building it. Redisson PRO emits that set through Micrometer, so it lands in whichever backend you already run — Prometheus, Datadog, CloudWatch, JMX and the rest of the Micrometer registry list. Provider classes live in org.redisson.config.metrics and attach to the config object directly.

Config config = ... // Redisson PRO config object

// JMX
JmxMeterRegistryProvider provider = new JmxMeterRegistryProvider();
provider.setDomain("appStats");
config.setMeterRegistryProvider(provider);

// Prometheus — wired through MeterRegistryWrapper, which adapts a registry you already have
PrometheusMeterRegistry registry = ...
config.setMeterRegistryProvider(new MeterRegistryWrapper(registry));

// Dynatrace
DynatraceMeterRegistryProvider p = new DynatraceMeterRegistryProvider();
p.setApiToken("");
p.setUri("https://.live.dynatrace.com/");
p.setDeviceId("myHost");
config.setMeterRegistryProvider(p);

Here are the metrics that map onto those four blind spots, plus per-node error rate and reachability. The first five rows sit under the per-node base name redisson.redis.<host>:<port>; the last is a per-object metric with its own base name:

Metric Answers

connections.active, connections.free, connections.max-pool-size

Is the pool exhausted? Free trending to zero is the alert, and it fires before the timeouts do

operations.latency

Latency as the client measures it, in milliseconds, as a histogram — including everything the slow log excludes

connections.reconnected, operations.retry-attempt

Reconnect and retry storms, which are otherwise inferred from a latency graph after the fact

operations.total-failed, operations.total-successful

Error rate per node, attributable to a specific server rather than to the cluster

status, type

Per-node reachability and role from this client's point of view. status is 1 connected, -1 disconnected; type is 1 MASTER, 2 SLAVE, 3 SENTINEL, as the docs label them

local-cache.hits, local-cache.misses, local-cache.evictions, local-cache.size

Near-cache effectiveness — the traffic Redis never sees. Reported per object, under base names such as redisson.local-cached-map.<name>

Per-object metrics follow the same shape, so an RMap reports hits, misses, puts and removals under redisson.map.<name>, and a topic reports messages-sent and messages-received. Distributed tracing is available separately through BraveTracingProvider and OtelTracingProvider for OpenZipkin Brave and OpenTelemetry, covered in Redis client tracing in Java. Both metrics and tracing are Redisson PRO features; the full metric list and per-provider configuration are in the observability documentation.

The point is not that client-side metrics replace server-side monitoring. It is that the two answer different questions, and an incident review that only has one of them tends to end with nobody able to explain what happened.

Redis Monitoring: Frequently Asked Questions

How Can I Monitor Redis Activity?

For live inspection, redis-cli MONITOR streams every command the server processes, though it is costly enough that it must be bounded and never left attached — see below. For ongoing monitoring, poll INFO — which is what every monitoring agent does under the hood — or run an exporter such as redis_exporter and scrape it with Prometheus. For diagnosing slowness specifically, SLOWLOG GET shows slow-executing commands and LATENCY LATEST shows latency spikes by event, though the latency monitor has to be enabled first because it is off by default.

What Is Redis Monitor Mode and How Is It Used?

MONITOR is a debugging command that streams back every command processed by the server, in real time, to any client that issues it. It is used to see what an application is actually sending, which is often not what its developers believe it is sending. Redis's own benchmark shows a single MONITOR client cutting SET throughput from roughly 95,000 to roughly 42,000 operations per second, and the documentation notes that more clients reduce it further. Administrative commands and QUIT are never logged, and AUTH appears with its arguments redacted. Use it briefly and never leave it attached. Note what it does not show you: run against a replica, it reports the traffic clients send to that replica interleaved with the replication stream from the primary, so it is no substitute for watching the primary. Reads served by the primary do not enter that stream — with the curious exception of PFCOUNT, which is flagged read-only but propagates when it rewrites its cached cardinality — and the primary's non-deterministic writes arrive already rewritten into their effects.

What Are the Best Redis Monitoring Tools?

There is no single best one, because they solve different problems. Redis Insight is the best free interactive tool for inspecting a running instance. Prometheus with redis_exporter, visualized in Grafana, is the standard open-source stack for continuous monitoring and alerting. Datadog, Dynatrace and New Relic are worth their cost mainly if you are already using them, since the value is correlating Redis against the rest of your stack rather than anything Redis-specific. On managed services the provider's own — CloudWatch for ElastiCache, Azure Monitor for Azure Cache — is the zero-effort baseline, though you can still point an exporter or agent at the endpoint from a host you control. Nearly all of them are ultimately parsing the same INFO output — AWS and Microsoft both document their managed metrics as being derived from it — so the differences are mostly operational rather than a matter of which metrics you can see. The providers add a few host-level figures Redis cannot report about itself, such as instance CPU and network utilisation.

How Do I Check Redis Memory Usage?

INFO memory gives the instance-wide picture: used_memory for what Redis has allocated, used_memory_rss for what the operating system sees, and maxmemory for the configured ceiling. MEMORY STATS breaks that down further and is more useful than most people realise, since keys.bytes-per-key and dataset.bytes separate your actual data from overhead. For a single key, MEMORY USAGE reports the bytes that key and its value require including administrative overhead, but note it samples nested values rather than measuring them — the default is five samples, so for large hashes, sets and sorted sets the number is an estimate. Pass SAMPLES 0 to count everything at O(N) cost.

What Is a Good Cache Hit Ratio for Redis?

Hit ratio is keyspace_hits divided by the sum of keyspace_hits and keyspace_misses, both from INFO stats. There is no universal target, because the right number depends on what you are caching. As rules of thumb rather than documented figures, a session store should sit near 100%, while a cache in front of a long-tail catalogue may be perfectly healthy at 70%. Redis Software's own observability documentation alerts below 50%, which is a reasonable floor for a general-purpose cache. Watch the trend rather than the absolute value, and read it alongside evicted_keys — a hit ratio falling while evictions rise is a memory sizing problem, whereas one falling with evictions flat is usually an application access-pattern change.

Why Is My Redis Slow When SLOWLOG Is Empty?

Because the slow log measures only command execution. The configuration documentation states that it excludes I/O operations such as talking with the client and sending the reply. Note this is narrower than usually reported: building the reply is counted, and only the socket write is excluded. The gap is one of cost per byte. Assembling a reply in memory is far cheaper per byte than pushing it through a socket, so a command can build several megabytes in a few hundred microseconds, stay under the 10 ms default threshold, and still take milliseconds to reach the client. Three things to check instead: reply size, which Valkey 8.1 exposes directly through COMMANDLOG LARGE-REPLY; host-level latency via redis-cli --intrinsic-latency, which establishes the floor below which tuning Redis achieves nothing; and client-side latency measured in your own application, which is the only place the full round trip is visible.

How Do I Monitor Redis Connection Pool Exhaustion?

Not from the server: the pool lives in your application, so Redis reports a healthy, modest connected_clients while your threads queue for a connection. You have to measure it client-side, and the raw signals exist in more than one Java client. Jedis exposes per-node pool counters through commons-pool2. Lettuce records command latency through Micrometer. What no open-source Java client does is publish those as one coherent per-node metric set to an external monitoring system without you wiring it up yourself. Redisson PRO publishes that set through Micrometer, so it reaches whichever backend you already run, covering the things the server cannot see: connection pool state as active, free and max-pool-size counts; operations.latency as a histogram in milliseconds measured from the application; reconnection and retry-attempt counts; per-node success and failure totals; and local cache hits, misses, evictions and size for near-cached objects. Per-object metrics cover maps, caches, topics, buckets, executor services and JCache. Distributed tracing is available separately through OpenZipkin Brave and OpenTelemetry.