Skip to content

Counting visitors

A visitor is one browser on one day. Rainlytics counts them from the viewer address CloudFront records, hashed under a salt that changes every day.

"visitors": { "distinct": 317, "additive": false }

That field rides on a rollup summary alongside the rows, on the questions that count pages. additive: false is there because two days of visitor counts do not add up, and the rest of this page is why.

The count is over the pageviews the summary’s question covers. A summary narrowed to /blog/ reports the visitors to /blog/, and one narrowed to a host reports that host’s. Views and visitors are two numbers over one set of rows.

Automated traffic is left out, the way it is everywhere else. Bots are most of a quiet site’s traffic and each crawler would otherwise arrive as a visitor a day.

One browser on one day, in one place.

  • Two devices are two visitors. A phone on the train and a laptop at a desk carry two addresses.
  • A household is one visitor, where two people behind one router run the same browser. The user agent goes into the hash for this reason and separates a phone from a laptop on the same address, and it cannot separate two identical iPhones.
  • A mobile carrier is fewer visitors than it should be. Carrier-grade NAT puts thousands of people behind one address, and the user agent is what splits them, imperfectly.
  • A VPN moves somebody, and the same person on and off one is two visitors.
  • A record with no address is nobody. CloudFront started recording addresses for Rainlytics in #73. Every day before that delivery change counts zero visitors.

So it is a measure of browsers rather than of people, and the number moves with how a site’s readers connect. Every privacy-preserving analytics product reports the same measure, and each of them counts a slightly different set of browsers. Compare the number against itself over time.

The salt changes at midnight UTC. The same browser carries one identifier today and a different one tomorrow.

A day of them counts. Two days added together count everybody who came back twice over, and a month of them is a figure about nothing. Thirty summaries each carrying "distinct": 429 are thirty numbers jq will happily sum, and the total describes nobody.

That is what additive: false says in the document, and what the VisitorCount wrapper says to TypeScript. A month of visitors is a query over raw, under the salt that month was counted with.

Hours work differently. Every hour of a day shares that day’s salt, and the hourly summaries of a day count the same identifiers the daily summary counts. They still fail to add, because somebody who came back after lunch appears in two of them. The daily summary is the answer for a day.

to_hex(sha256(to_utf8(concat(<the day's salt>, '|', c_ip, '|', cs_user_agent))))

Athena computes it while counting, and every digest dies with the query that made it. The summary holds the count and no more.

The three parts are joined by |, which an address cannot contain. The text hashed for one address and user agent therefore belongs to that pair alone.

SHA-256 rather than the faster xxhash64. Both are in Athena engine version 3, and a 64-bit non-cryptographic digest is forgeable by anybody holding one. At the volumes a site of this size produces, the speed makes no difference worth having.

The salt for a day is derived from one secret and the date:

salt(day) = HMAC-SHA256(secret, "rainlytics/visitor-salt/1/" + day)

The secret is a SecureString in SSM Parameter Store. The Lambda that computes the summaries reads it once per run and derives the salt for each window it is computing. Four things follow, and they are the four the decision in #53 asked for.

Every record of a day counts under one salt. The day comes from the window, and every window inside a day gives the same date.

Tomorrow is a different salt. The date is in the message the HMAC is taken over.

A re-run of a day reproduces it. The date comes from the window being computed and never from the clock. A window recomputed next week therefore writes the count that was there before it. The summary schedule recomputes a trailing window on every run for exactly this, and a salt taken from the clock would make the second run disagree with the first.

A reader of the log bucket cannot get it. The secret lives in Parameter Store alone, away from the bucket, the summaries, the CloudFormation template and the schedule that carries the query. It is encrypted at rest under the aws/ssm managed key, and reading it takes ssm:GetParameter on that one parameter.

The secret is meant to stand rather than rotate. Replacing it makes every day from then on count somebody new, and makes every day before it uncountable. The date is what rotates.

Athena takes no secret of its own. The salt goes into the statement as a quoted literal, and a copy of it then lives wherever Athena keeps its query history (45 days, behind athena:GetQueryExecution on the workgroup). CloudTrail records the query string as ***OMITTED*** for StartQueryExecution, and the statement reaches S3 nowhere.

This is why the statement carries a day’s salt and never the secret. HMAC is built so that a key cannot be recovered from a message and its digest. A salt read out of query history therefore covers the days it appears for, and says nothing about any other day or about the secret.

Nothing creates it for you. CloudFormation writes String and StringList parameters and no SecureString, and a construct that generated a secret at synthesis would put it in a template, which is the one place it must not be.

Terminal window
aws ssm put-parameter --name /rainlytics/visitor-salt --type SecureString --value "$(openssl rand -hex 32)"

Run it once per deployment, in the account and region the summaries run in. The summary schedule construct takes visitorSaltParameter for a name of your own, and grants the job ssm:GetParameter on whichever one it was given.

A run that meets no parameter fails and says so, naming the parameter and printing that command. Only the questions that count visitors read it. A deployment that counts none needs no parameter and never asks for one.

pageviews alone. A rollup says so with countsVisitors:

import { pageviews, type Rollup } from "@kensio/rainlytics";
const blogVisitors: Rollup = {
...pageviews,
name: "blog-pageviews",
countsVisitors: true,
};

The count is always over pages, whatever the question beside it counts. A summary of status codes carrying one would report a number about rows it never looked at, and the field is left out of every question that counts something else. Absent and { "distinct": 0 } mean different things, and a reader can tell them apart.

It costs one extra Athena query per window per run. The five shipped questions on both cadences, recomputing two windows, come to 250 queries a day and about 38 cents a month. The visitor count on pageviews adds 50 of those, which is about 8 cents. The summary schedule page has the arithmetic.

The raw log bucket holds viewer addresses in the clear, for as long as it holds anything. That is the price #53 paid for a visitor count, and the log bucket page has the expiry that decides how long it lasts.

The salt protects the identifier and never the source. Anyone who can read the log bucket has the addresses themselves, at better resolution than any digest would give them.