COLITU LAB14 of 21 probes online

Research / App test report 001

Does the connection check really go through the tunnel?

An outside reviewer suggested a test: let the tunnel handshake succeed but drop everything inside it, and see whether the app still calls the connection verified. We ran it on the Windows and Linux apps, during fallback and across a network change, and checked app traffic and DNS too. It found three problems; all are fixed.

Tested 10 Oct 20266 min read

The suggestion

After reading how Colitu’s apps decide that a connection works, an outside reviewer proposed a route-provenance test. The idea: make the probe endpoint reachable over the device’s ordinary route, let the tunnel handshake succeed, but drop every packet inside the tunnel. A correct verifier must then report failure, even though an ordinary request outside the tunnel would succeed. They asked for the same test during a fallback from one transport to another and after a network change, and for application traffic and DNS to be checked as well as the probe. They had not run Colitu and did not claim it leaks; it was a suggested test for the state machine. We ran it on 10 October 2026.

Does the verifier bind its request to the selected tunnel?Now always. Before Windows 2.8.2 and Linux 1.4.2, only partly.

The connect check with a warm spare was bound to the primary server. The watcher, the check after sleep, the restart check and the connect check without a spare used the ordinary local proxy port, where routing rules apply.

Is an old success invalidated when routing changes?Now yes on Windows (2.8.4). Before, only when the network adapter changed.

Joining another Wi-Fi network on the same adapter left the old verification standing for about 85 seconds. The Linux app already re-checked about 4 seconds after any address change.

Test 1: handshake works, tunnel is a blackhole

We took the Windows app’s real generated core configuration (sing-box, with a warm spare and every routing rule) and replaced only the two server outbounds with local VLESS servers. A blackhole server accepts the handshake and drops everything sent through it; a healthy server forwards normally. The probe endpoints stay reachable directly from the machine. The check is the app’s own: three generic generate_204-style endpoints at once, the first 2xx answer wins, two rounds of 7 seconds.

Scenario 4 adds one routing rule that sends the probe hosts direct. That is what split tunnelling in “only selected apps use the VPN” mode does for everything outside the selected apps.

ScenarioPrimary onlySpare onlyOld check pathNew check path
1. Both servers healthy (control)verifiedverifiedverifiedverified
2. Primary blackholed, spare healthy (fallback A → B)fails, correctverifiedverified via spareverified via spare
3. Both blackholedfailsfailsfailsfails
4. Both blackholed, probe hosts routed directfails, correctfailsVERIFIED, wrongfails, correct

Scenario 4 is exactly the failure the reviewer described: a dead tunnel reported as working because the check took another route. The fix gives every configuration a loopback check inbound, colitu-check, whose first routing rule goes to the tunnel itself (the server, or the group of primary and spare). All checks of the whole tunnel now go through it, so no routing rule can send them anywhere else. The behaviour is covered by unit tests on both cores, with and without a spare, and by re-running this fixture.

Test 2: application traffic and DNS on a real device

On a second Windows 11 machine, a script sampled every 2 seconds: the route Windows picks for the internet, the exit address of an ordinary HTTPS request, which resolver a normal system DNS lookup reaches (a TXT lookup that returns the resolver’s address), and whether a request bound to the physical network adapter, outside the tunnel, gets out.

Windows app, connection pathKill switchApp trafficDNS lookupsRequest outside the tunnel
2.8.2, Xray-based transportofftunnelhome ISP resolvergets out (expected)
2.8.2, Xray-based transportontunnelno answerblocked
2.8.3, any transportontunneltunnel, ~550 msblocked
2.8.3, any transportofftunneltunnelgets out (expected)

The cause: on Windows, connections carried by an Xray-based transport used Xray’s own TUN adapter, and that adapter gets no DNS server. Windows kept asking the resolver of the Wi-Fi adapter. The traffic itself went through the tunnel; the names did not. With the kill switch on, those lookups were blocked, so names did not resolve at all. Since Windows 2.8.3, sing-box runs the TUN adapter for every transport and hands those connections to Xray, as the Linux app already did. The Linux app was not affected.

Test 3: network change on the same adapter

Same machine and sampling, kill switch on. The machine moved from home Wi-Fi to a phone hotspot and back, always on the same Wi-Fi adapter. Times are relative to joining the hotspot.

Windows 2.8.3

  1. Joined the hotspot. The app keeps showing “protected”.
  2. System DNS lookups start to time out: the tunnel’s DNS connection was opened on the old network.
  3. The connection watcher misses its first check. The next one comes 30 seconds later.
  4. Second miss. Recovery would have started at the third, around t+85 s.
  5. Every request outside the tunnel was blocked. The only sample that got out was taken while the user turned the VPN off by hand; the kill switch releases on purpose then.

Windows 2.8.4

  1. “Network changed on Wi-Fi (new address or gateway)”: the app shows “reconnecting”, restarts the tunnel on the new network and verifies it again. The kill switch stays armed.
  2. The restarted transport could not be verified on this mobile network, so the app moved on to a full reconnect. In the released build this restart check is one 5-second round instead of two 7-second rounds.
  3. Back on home Wi-Fi: no interruption seen, every sample inside the tunnel.
Still open. After a network change, the reconnect orders transports by what worked on the previous network. On this mobile network it first tried two transports the network does not let through, 6 seconds each, before the one that works. A fix is in progress.

What changed

VersionChange
Windows 2.8.2 · Linux 1.4.2Every tunnel check goes through the colitu-check inbound; routing rules can no longer send a check around a dead tunnel.
Windows 2.8.3sing-box runs the TUN adapter for every transport. DNS stays inside the tunnel with the kill switch off, and resolves with it on.
Windows 2.8.3 · Linux 1.4.3The warm spare is attached on Xray-based transports again. A configuration check had refused it since 2.8.0 / 1.4.0 over a leftover routing rule that can never match.
Windows 2.8.4A new network on the same adapter invalidates the last verification, shows “reconnecting” and rebuilds the tunnel within seconds.

Limitations

  • Windows was tested on real devices. Linux was covered by unit tests and the same fixture logic, not on a Linux machine. Android, iOS and the browser extension use different network stacks and were not part of this test.
  • The blackhole fixture used local VLESS servers for both paths. A blackholed UDP transport (Hysteria2) was not modelled.
  • The network change on 2.8.4 was run with the kill switch on. The kill-switch-off case was measured in steady state on 2.8.3 only.
  • Samples were 2 to 13 seconds apart, so a single packet escaping between samples would not show. The request bound to the physical adapter is a strong indicator, not a packet capture.
  • We do not name the networks, operators or servers, and we publish no addresses.

Thanks to the reviewer who suggested the test. Release notes: Windows, Linux. Source: colitu/windows, colitu/linux.