COLITU LABProbes starting

Methodology

How Colitu Lab measures

What a test is, how a result becomes a rate and a status, how events and fingerprints are derived, what the anonymous app data are, and what is never published. Everything here describes the code that produces the numbers; when the code changes, this page changes with a new version.

Methodology v1.0 · 6 October 2026Data & API

Data sources

The lab has two sources, kept apart everywhere they appear:

  • Probes — vantage points that test every transport to every public Colitu server on a schedule. Each result is stored and published individually, as measured. They are the lab's primary source.
  • App data — the transport results Colitu's apps already report while connecting. The lab keeps only hourly counts per country and network, and publishes them late and only where enough devices stand behind them (see App data). Pages label them “App data (delayed)”.

In addition, the team can record notes: findings written by hand, shown among the events with the tag “Note”. Notes are never produced automatically and never counted in a rate.

Probes

A probe is a program (labprobe) that runs on a host and reports to the lab. Every probe has a short name, a location (country and city), a network type — datacenter, residential or mobile — and the operator of the host. Today every probe runs on a host Colitu operates; the probes page lists them.

  • Server list. The probe fetches the servers the way a VPN app does: from the subscription of a dedicated lab account. It therefore tests what users are offered, with the same settings.
  • Schedule. One round every 20 minutes. Results that cannot be delivered are kept on disk and sent with the next report; the lab accepts results up to 6 hours old.
  • Network. The probe's network (its ASN and the network's name) is looked up from the address its reports arrive from. The address itself is not stored.
  • Online. A probe counts as online when it has reported within the last 90 minutes. “Reporting, 7 d” is the share of hours in the last 7 days (or since the probe was added) with at least one stored result.

Tests

A round has three kinds of test. Each has a time limit of 12 seconds.

1. Baselines — the network without a VPN

  • DNS over UDP and DNS over TCP: a lookup of example.com sent directly to the public resolvers 1.1.1.1 and 8.8.8.8.
  • HTTPS to three neutral sites: wikipedia.org, google.com and cloudflare.com. A response with a status below 500 counts as success.

If not a single baseline succeeds, the probe's own connection is down: the round reports the baselines and skips the tunnel tests, so an outage at the probe does not show up as failing VPN transports.

2. Reach — can the server's port be reached at all?

A plain TCP connection to the transport's port on the server, closed immediately. Only transports that run over TCP have a reach test; Hysteria2 and TUIC use QUIC over UDP, which has no equivalent connection to open.

3. Tunnel — does traffic actually flow?

The probe starts the real client core for the transport — Xray for VLESS, Trojan and Shadowsocks, the Hysteria client for Hysteria2 — with the settings from the subscription, waits up to 5 seconds for its local SOCKS port, and requests https://www.google.com/generate_204 through it. Every request uses a new connection. A response with a status below 500 is a success; the time recorded is from sending the request to receiving the response (“time to first answer”), so it includes the tunnel's handshake.

After a successful tunnel test the probe downloads a 2 MB file from speed.cloudflare.com through the tunnel for at most 10 seconds, and records the throughput if at least 64 KB arrived. It is a rough indication, not a speed test: the probe's own line, the server's load and the route all affect it.

Comparing reach and tunnel on the same port is the core of the method. A port that answers while the tunnel through it fails points at inspection of the traffic; a port that does not answer points at blocking of the address — or at a server that is down, which the other vantage points then see too.

Times

Medians are taken over successful tests only: tunnel time for tunnel tests, connection time for reach tests and baselines.

Failures

A failed test carries one failure class, derived from the error the probe saw:

CodeShown asMeaning
timeoutTimed outNo answer within the time limit (12 s per test).
resetConnection resetThe connection was reset, by the far end or by something on the path.
refusedRefusedThe port actively refused the connection.
unreachableUnreachableNo route to the host or network.
tlsTLS errorThe TLS handshake or certificate check failed.
dnsDNS failureA name could not be resolved.
eofClosed earlyThe connection was closed before an answer arrived.
otherFailedAnything else, including an HTTP error (status 500 or above) through the tunnel.
probe_errorProbe errorA problem on the probe itself, such as a client core that did not start. Left out of every rate.
authCredentials refusedThe server refused the lab account's credentials — a problem on our side. Left out of every rate.

probe_error and auth say nothing about the network, so they are stored and published in the raw data but excluded from every rate, median, status, event and fingerprint. Every other failure counts. Failure classes are a hint, not a diagnosis: a filter can make a connection time out or reset it, and so can an overloaded server.

Rates and status

The success rate is successful tests divided by all tests in the selection (excluding the two classes above). Unless a page says otherwise it is pooled over every server, and for a country over every vantage point in it. The status turns the rate into a five-step scale:

StatusSuccess rate
Available98–100 %
Degraded90–98 %
Restricted50–90 %
Heavily restrictedbelow 50 %
Not enough datafewer than 6 tests

A selection with fewer than 6 tests is “Not enough data”, whatever its rate. Changes next to a rate are in percentage points against the window of the same length just before; they appear only when that window has at least 6 tests.

Charts group tests into buckets — 1 hour for windows up to 48 hours, 2 hours up to 8 days, 6 hours up to 31 days, and 1 day beyond — and leave a gap where a bucket has no tests. A gap means “not measured”, never “0 %”.

Events

Every 10 minutes a detector looks at each probe and transport separately, with all servers pooled. It compares the last 2 hours with the 7 days before:

  • It needs at least 8 tests in the last 2 hours and 30 in the week before; otherwise it does nothing.
  • An event opens when the week's rate was at least 60 % and the recent rate is at least 25 percentage points lower. It is an outage if the recent rate is 10 % or less, otherwise a degradation. Its start is the beginning of the 2-hour window.
  • While an event runs it is judged against the rate before it began, keeps its worst reading, and becomes an outage if the rate falls to 10 % or less.
  • It closes when the recent rate is back within 10 percentage points of that earlier rate.

Because servers are pooled, one server going down is usually too small a share of the tests to open an event; the per-server-location tables on the protocol pages show that case. An event says that a transport stopped working from a vantage point, not why: the baselines and the fingerprint help tell interference from an outage.

Network fingerprint

The network pages read five behaviours off a network's results in the selected window. Each compares two kinds of test from the same network, so that a server outage — which fails both — does not look like interference. Negative differences count as zero.

SignalValueObserved atNot observed at
Server addresses unreachable (IP or port blocking)Failure rate of reach tests≥ 20 %≤ 5 %
TCP transports fail although the port answersTunnel failure rate of the TCP transports minus their reach failure rate≥ 15 points≤ 3 points
QUIC/UDP transports fail more often than TCP onesTunnel failure rate of Hysteria2 and TUIC minus that of the TCP transports≥ 15 points≤ 3 points
Plain DNS over UDP to public resolvers disruptedFailure rate of DNS over UDP minus DNS over TCP, same resolvers≥ 15 points≤ 3 points
HTTPS to neutral websites fails without a VPNFailure rate of the HTTPS baselines≥ 10 %≤ 2 %

A signal needs at least 12 tests behind it (for a difference, 12 on the smaller side); with fewer, or with a value between the two thresholds, it is inconclusive. The thresholds are deliberately coarse. “Observed” means the measurements are consistent with that behaviour, not that a particular filter has been identified.

App data

Collection of app data is currently off. It stays off until it is announced here and in the privacy policy; until then every figure on this site comes from the probes.

When a Colitu app connects, it tells the service which transports it tried on which server and whether traffic flowed — the apps use this to switch transports automatically. When app data are on, the lab turns these reports into counts:

  • Address. The country and network (ASN) of the address a report comes from are looked up; the address is then discarded and never stored.
  • Counts. The lab adds the report to hourly counts per country, network, transport and server country: attempts, successes and the sum of connection times (pages show the mean).
  • Devices. To count how many different devices stand behind a cell, the device identifier is turned into a pseudonym with HMAC-SHA256 and a random key created for that UTC day. The identifier itself is not stored. At the day's roll-up the pseudonyms are counted, then deleted together with the day's key, so pseudonyms of different days cannot be linked.
  • Publication threshold. A country's day is published only if at least 5 different devices contributed that day (for a network page, at least 5 on that network).
  • Embargo. Only days that ended at least 7 days ago are published, so app data never say what works where right now. Pages show the 30 days before the embargo; the field-daily export goes back up to a year.
  • Through the tunnel. Reports sent from a Colitu server's own address describe no access network and are ignored, as are reports about servers in locations the site does not list.
  • Opting out. An app request carrying the header X-Colitu-Lab: off is never counted. A switch for this in the apps' settings is planned.

Selection bias. Apps try the next transport only after one fails, so transports early in the order are tried far more often, and later ones mostly on networks where the earlier ones already failed. App-data rates are therefore not comparable between transports the way probe rates are; they are useful for seeing changes within one transport and country over time.

Coverage and bias

  • Vantage points. Currently 0 probes in 0 countries, all on hosts Colitu operates. A datacenter connection is often filtered less than a home or mobile one, so a datacenter probe can miss interference that targets consumer networks. Each result carries the probe's network type.
  • Servers. Only Colitu's public servers are tested, and only the transports they offer (TUIC is not offered at present). A result describes a transport on these servers from that network, not the transport everywhere.
  • One target. The tunnel test requests one small URL. It shows whether a tunnel carries traffic, not whether every website is reachable through it, and it is not a measure of web censorship in general.
  • Countries without a probe appear only through app data, and only when those are on and above the thresholds.

Conflict of interest

Colitu runs both the probes and the servers they test, and offers a VPN that uses these transports. We say so plainly because it matters for how the results should be read:

  • The results describe transports on these servers, from these networks. They are not a ranking of VPN providers, and no other provider is measured.
  • No transport gets special treatment. Every transport is tested the same way, in the same round, with the settings users receive, and every failure counts except the two classes that are our own fault (probe errors and refused credentials) — which are still published.
  • Everything is published raw: every probe result, the events, the daily aggregates and this method. Anyone can recompute every figure on this site from the data and check it.

What is published — and what never is

Published: every stored probe result — time, probe, the probe's country, network and network type, test, transport, the country of the server, the target of a baseline, the result, the failure class, the times and the download throughput — plus events, daily aggregates and app-data counts that pass the thresholds.

Never published:

  • server identities and addresses (names, IP addresses, ports, keys) — only the server's country;
  • the IP addresses of the probes' hosts;
  • users' IP addresses or device identifiers, which the lab never stores;
  • app data in real time, or for any cell below the device threshold.

Tests that reached a server in a location the site does not list are not stored at all, for probes and app data alike, so they cannot leak through an export.

Retention

  • Probe measurements and the hourly app-data counts: 2 years, then deleted.
  • Daily device pseudonyms and their keys: deleted at the day's roll-up, so a pseudonym exists for at most about one day; a separate clean-up deletes any pseudonym older than 2 days as a safety net.
  • Users' IP addresses and device identifiers, and the probes' addresses: not stored.

Licence and citation

The data and the charts are licensed under Creative Commons Attribution 4.0 (CC BY 4.0). Use them freely, including commercially, and cite “Colitu Lab, lab.colitu.com”. The Data & API page has a citation and BibTeX entry. Network names come from DB-IP ASN Lite, also under CC BY 4.0. The map uses Natural Earth.

Changes

VersionDateChange
1.06 October 2026First published methodology.

Questions or corrections: hello@colitu.com. The settings in force right now (thresholds, embargo) are also in the API: overview, field field.