Why a fault lab
Adaptive Connect is the part of the Colitu apps that picks a server and a transport, watches the tunnel and moves to another one when something breaks. It is easy to say it “switches automatically”; it is harder to say how long a user is without a connection when a protocol is blocked, a server disappears or a network starts dropping packets. We wanted a test that answers that with numbers, runs the same way every time, and compares one app version with the next.
How it works
The tool inserts firewall rules on our own servers that match only the test device’s public address and carry a tag; every rule has a timer on the server that removes it even if the tool dies. No other user is affected. Every 5 seconds the phone makes a real HTTP request through the tunnel. For each fault we record the longest stretch without traffic (inside the fault and the 60 seconds after it), the first good request after it began, and how long traffic took to return after it was lifted.
| Fault | What the rule does | What it imitates |
|---|---|---|
| Protocol drop (Hysteria2, VLESS Reality, Trojan, VLESS XHTTP) | drops that protocol’s ports on every server | a network that blocks one protocol |
| Reset | answers VLESS Reality connections with a TCP reset | active connection killing |
| 16 KB freeze | drops every TCP flow after 16 KB | the “first kilobytes, then silence” pattern seen on Russian networks |
| UDP block | drops all UDP to every server | networks that throttle or block UDP and QUIC |
| Loss 5 % / 20 % | drops that share of packets at random | a bad mobile or Wi-Fi link |
| Server in use | drops every port of the server carrying the tunnel at that moment | a server that goes down or gets blocked |
The suite and the comparison tool are plain scripts. The raw results of both runs below are available as JSON without addresses: baseline (2.8.0) and test build.
Run 1: the released app
Android 2.8.0 with Adaptive Connect 2.0, automatic mode, on a home Wi-Fi in Russia. On this network every TCP-based transport stops carrying data 30–50 seconds after it connects, even without any fault from us; only Hysteria2 (UDP) keeps working.
| Fault | Longest outage | Back after the fault ended | Reading |
|---|---|---|---|
| Hysteria2 dropped everywhere | 95 s | 2 s | no working protocol left on this network |
| VLESS Reality dropped | 25 s | 3 s | was the transport in use |
| Trojan, reset, 16 KB freeze, loss 5 % and 20 %, XHTTP | 0 s | 1–5 s | |
| UDP blocked | 100 s | 31 s | slow return, see below |
| Server in use cut | 45 s | 3 s | warm spare was dead, see below |
397 of 466 probes succeeded (85.2 %). The outages during the Hysteria2 drop and the UDP block are expected on this network: while they last, no transport works. The interesting numbers are what happened around them.
What the logs showed
When Hysteria2 stalled mid-session, it went to the back of the line on that server for 10 minutes. The app then cycled through TCP transports that freeze on this network, and only returned to Hysteria2 once every transport carried a stall mark. A second detail made it slower: the miss counter was reset before the 60-second gap between automatic switches was checked, so a dead tunnel waited out the gap and then needed three more misses.
Replacing a spare reloads the core for about a second, so the app waits for an idle moment: under 10 KB of traffic in 10 seconds. On a real phone, background sync plus the app’s own checks measured 15–30 KB per 10 seconds, all the time. The swap was deferred indefinitely, so the spare was dead when it was needed.
What we changed
- A transport that has already worked on the current network gets a 90-second mid-session penalty instead of 10 minutes; others keep 10 minutes.
- Misses keep counting inside the gap between automatic switches, so the switch happens as soon as the gap ends.
- After waiting two minutes for an idle moment, light traffic (under 32 KB per 10 s, below a voice call) no longer blocks replacing a dead spare.
- Separately, from another test on Windows: when the device joins a different network, the memory of the old one (stalls, last good transport) is no longer applied to the new one.
These are in a test build for Android and in the Windows and Linux sources, with unit tests; they are not released yet.
Run 2: the test build, and why it proves less than it seems
| Fault | Released app | Test build |
|---|---|---|
| Hysteria2 dropped everywhere | 95 s | 0 s |
| VLESS Reality dropped | 25 s | 0 s |
| UDP blocked | 100 s | 0 s |
| Server in use cut | 45 s | 35 s |
| All other faults | 0 s | 0 s |
| Probes that succeeded | 85.2 % | 98.5 % |
Do not read this table as “the fix removed the outages”. In the second run the phone had settled on a different server, reached over a TCP transport that does not freeze there. With TCP working, a Hysteria2 drop or a UDP block simply does not touch it. The situation the first fix targets (Hysteria2 in use, UDP blocked, then released) did not occur, so that fix is not yet verified on the device. When the server in use was cut, the log shows the spare was again dead with nothing on offer to replace it, so the spare fix did not come into play either. The run shows no regression; it does not show the improvement.
What we learned about the lab itself
- The first attempt listed eight servers. Halfway through, the phone moved to a server outside the list and every later fault missed it. Since then all servers the device could pick are in scope, and the tool logs which server carries the tunnel before each fault.
- The rules match an address, and a home has one address for every device behind it. Other computers in the same home lost their connection during the faults. Runs are now announced in advance, and the workstation running the tool is moved to a server left out of the test.
- Rules on all servers are applied in parallel so a fault starts and ends at the same moment everywhere; a hung server no longer stops the run, and an interrupted run can resume at a given fault.
Limitations
- One Android phone, one home network in Russia, one day. The numbers describe that combination, not Colitu in general.
- The two runs did not start from the same state (server and transport in use), so they are not a clean before-and-after comparison.
- One fault in the second run (cutting the server in use a second time) was skipped: the tool did not see the tunnel on a listed server in its short capture window.
- Delay, jitter and DNS tampering are not covered; they need shaping on the device side.
- The Windows and Linux changes are covered by unit tests only, not by this lab.
- Next: a short targeted run that puts the phone on Hysteria2 first and then blocks UDP, to verify the first fix on the device.