zeroreads Open source log reduction, with proof

Cut the log lines nobody reads.

zeroreads reads every query, alert and dashboard that touches your logs, proves which lines have zero reads, and writes the pipeline config that removes, shrinks or archives them.

Open source under Apache-2.0. Runs on your own infrastructure and sends nothing anywhere.

Read the quickstart What has been measured

1

Logs arrive faster than anyone reads them.

Most of the volume is a handful of repeating templates, such as housekeeping, heartbeats and retries that worked. The lines pressed into this sheet are templates from the OpenTelemetry demo, twelve services with their own load generator, run on a local kind cluster.

2

Every reader leaves a mark.

zeroreads reads the Loki query log, ruler alerts, every Grafana dashboard, alert and stored query in every organisation, and OpenSearch in beta. Each query is checked against the exact pattern of each rule, and a query it cannot model counts as reading everything. Each lamp here is one of the demo's readers, and red pencil marks what it read.

3

Zero reads, proven for each rule.

A rule acts only when no logged query, alert, dashboard or stored query can return one of its lines. Every reader in the report either shows a line it reads or names what it had to assume. Missing evidence blocks every rule until it is fixed or deliberately acknowledged. Struck through in lead, the templates no lamp reached.

4

Measure before anything is removed.

Shadow mode removes nothing. It counts, rule by rule, what enforcing would remove. In the demo the Collector counted 383 lines where Loki stored 382, the window edges differing by seconds.

5

Then shrink, archive or drop.

Aggregate turns lines into a count, dedupe collapses repeats, sample keeps an exact share, archive moves lines to storage you choose, and drop is there when your policy allows it. Warnings and errors are never touched. Here the struck lines are pressed flat into counts, which is what aggregate does.

6

verify keeps watching.

On a schedule, verify reads every source again. When someone adds a dashboard that reads a removed line, here one that reads Sending Quote, the job fails with exit code 3, says which reader, and prints the rules that are still safe to deploy as the revert.

It writes config for the pipeline you already run

Runtimes zeroreads writes configuration for, and the versions it is checked against
RuntimeWhat it writesChecked against
OpenTelemetry Collectorfilter, logdedup, transform, signal_to_metrics, forwardcontrib 0.161.0, full loop on Kubernetes
Vectorremap, route, reduce, log_to_metric0.58.0, every event compared
Fluent Bitmodify, grep, lua, log_to_metrics5.1.2, every record compared
Telemetry Policypolicy files for policy-go runtimespolicy-go 1.12.1, only with an explicit flag

Evidence comes from Loki 3.7.8, Grafana 13.2.2, and OpenSearch with Dashboards 3.8.0 in beta.

What has been measured

Every number comes from the test suites in the repository, run against the real engines above on a local kind cluster.

9,114
"nobody reads this" answers checked against what Loki returned, over random, fuzzed and adversarial queries. None was wrong, which bounds the error rate under 0.033% at 95% confidence. 144 OpenSearch answers were each confirmed by OpenSearch.
38%
of "this query reads it" answers are exact, and all 388 exact ones that were checked are confirmed: the line shown was stored in Loki and the query returned it. The rest name what they assumed, so over-blocking in an install can be seen and fixed.
2,365
Loki queries from 358 public dashboards. 90.3% parse, and Loki 3.7.8 rejects every one that does not. 52.5% pick their streams with template variables, which count as reading every service they could name.
16.1%
of the stored lines in the demo belonged to the 14 of 26 rules that acted, 27.3% of the bytes. While enforcing none of them was stored, reconcile matched all 14, and verify passed.
1m45s
to analyse 999,999 log lines and 300,000 query executions at 65 MB. A Grafana with 5,000 dashboards takes 2m33s at 338 MB.

A real report, from the demo run

The rules that acted, their exact templates and the lines measured in a ten minute shadow window.

ServiceTemplateLinesAction
kafka[ProducerStateManager partition=__cluster_metadata-0] Wrote producer snapshot at offset <*> with 0 producer ids in <*> ms.40aggregate
kafkaDeleted producer state snapshot <*>39aggregate
kafkaDeleted log <*>39aggregate
kafka[SnapshotEmitter id=1] Successfully wrote snapshot <*>31aggregate
checkout[PlaceOrder]11aggregate
shippingSending Quote10aggregate

One detail from the same run: an ad hoc query for |= "quote" read Requesting quote, so that template kept every line, while Sending Quote has a capital Q and no reader.

Limits, plainly

  • Only Loki is analysed line by line. OpenSearch is beta and decided per service. Rollups are experimental.
  • verify catches a new reader after the fact. It cannot bring back lines removed before the revert, which is what the archive action is for.
  • Value depends on dashboards selecting streams by your scope label. A query that selects by namespace or a template variable is assumed to read every service it could.
  • Services that log JSON rarely qualify, since a line filter on JSON can match any field.
  • No production cluster has been measured yet. The numbers above come from test suites and a demo app.

Install

Build it from source with Go 1.27.1, or take a signed release.

go install github.com/Bisman-Singh/zeroreads/cmd/zeroreads@latest

zeroreads analyze -c zeroreads.yaml -o out/
zeroreads emit -c zeroreads.yaml -rules out/rules.json -mode shadow -o collector.yaml
zeroreads verify -c zeroreads.yaml -rules out/rules.json -deployed collector.yaml

Run verify on a schedule with the Helm chart in charts/zeroreads, or in CI with the GitHub Action in the repository. Releases carry checksums, SBOMs and keyless signatures, and the README shows how to verify them.