Colitu Adaptive Connect 2.0 and Warm Spare: a technical report
Colitu Adaptive Connect 2.0 written up as a short research report. Server ranking, per-network memory, anonymous network hints, health check, warm spare and mid-session watching against networks that block or freeze connections, plus method, results and limitations.
*Technical report. This page explains the design, the measurement method and the results behind the behaviour described in How Colitu Adaptive Connect works and Warm spare. The behaviour described here is in the latest versions of the apps.*
Abstract
Some networks block VPN protocols; others behave more subtly: they let the connection come up and then quietly stop the traffic a few dozen seconds later. On such a network "answers the ping" and "carries traffic" are not the same thing. This report describes the set of mechanisms the Colitu apps use against this problem, Colitu Adaptive Connect 2.0: server ranking that takes proximity and privacy into account, per-network memory kept only on the device, network hints derived from anonymous counters, a connect-time health check based on real traffic, server fallback, a warm spare kept ready inside the running connection, and a mid-session watcher that covers every transport. The evaluation was carried out on a single Android phone on a single residential Wi-Fi network in a country with active deep packet inspection (DPI), with faults injected at the server for this one client only. In the best configuration, when the main path was cut for 120 seconds, the VPN stayed on, there was no reconnect and the outage lasted about 5–10 seconds. In a 60-minute soak test with 11 faults injected on 8 servers, 98.8 % of the probes through the tunnel succeeded; when the server in use was fully cut for 101 seconds, the user saw only a stall of about 5 seconds and the VPN stayed on. Limitations and situations not yet measured are listed separately.
1. Introduction
- The first job of a VPN app is to connect to the right server with the right connection mode. Colitu servers offer five connection modes from four protocol families: Hysteria2 (UDP/QUIC), VLESS Reality, VLESS XHTTP, Trojan and Shadowsocks 2022 (see Technical details).
- The classic approach ranks servers by ping and picks the fastest responder. But ping is only a hint: on some networks a server answers the ping and still carries no traffic at all.
- There is a harder case: the connection comes up, the app says "Connected", but after a while traffic stops while the TCP connection still looks open. For the user this means the VPN is on but the internet is gone.
- Adaptive Connect 1.x solved part of these problems, but our measurements showed two important gaps: the server list came in database-id order, and the mid-session watcher only covered Hysteria2 (see section 4).
- The goal of Adaptive Connect 2.0 is: find a path that really works when connecting, notice a path that breaks while connected and move the user to another path, if possible without turning the VPN off, and do all this without collecting information about the user.
2. Threat and failure model
The design considers the failure types below. They are based on behaviour seen on some networks and with some mobile operators, not on any particular country.
| Failure | What happens on the network | What the user sees |
|---|---|---|
| Protocol blocking | A given protocol's connection never comes up | Cannot connect |
| Silent freezing | The connection comes up, traffic stops after N seconds, TCP stays open | "Connected", but pages don't load |
| UDP blocking | UDP-based modes (Hysteria2) don't work, TCP does | Cannot connect in the UDP mode |
| Server outage | A server can't be reached with any protocol | No connection on that server |
| Network change | Switching between Wi-Fi and mobile data, Wi-Fi dropping | Short drop; different blocking on the other network |
Assumptions:
- A network may block a mode while another network doesn't, so "works" is a property of the network.
- A network may freeze all TCP modes of the same server together; in that case another TCP mode on the same server is useless as a spare.
- The device's own internet can go away; in that case no mode should be blamed.
3. Design
The mechanisms below were designed for the Windows, Linux, Android (including TV) and iOS apps. All mechanisms are the same in every app. The measurements in section 4 were made on one Android phone.
3.1 Server ranking
- A location the user picked is used as is. Colitu never switches a manual choice on its own.
- In automatic mode ("Fastest server") the order is:
| Rank | Criterion |
|---|---|
| 1 | The server that last worked on the network you are on now (remembered 24 hours) |
| 2 | Lowest measured ping; a ping older than 10 minutes or measured on another network doesn't count |
| 3 | Servers in other countries before servers in your own country, nearer countries first (also the order when there is no ping yet) |
| 4 | A server that didn't answer the ping goes to the end |
| 5 | A server that failed on this network in the last 30 minutes goes last |
- "Recommended" in the server list always shows the server automatic mode would connect to.
- Privacy: To order the list, the panel works out your country and your internet provider's network from the IP address of that request only. Nothing is stored. When the request comes through the VPN the panel can't know them, and the app keeps the last values it learned.
3.2 Per-network memory
- A "network" is the connection type (Wi-Fi, mobile data, Ethernet) combined with your internet provider's network. Home Wi-Fi and mobile data are remembered separately.
- The memory is kept only on the device:
| Entry | Lifetime |
|---|---|
| Last working server | 24 hours |
| Last working mode per server | 24 hours |
| A mode that carried no traffic ("stall mark") | 6 hours (for exceptions see 3.8) |
| A server where everything failed | 30 minutes |
- Old entries are deleted; at most 200 entries are kept.
- As a result, a mode blocked on mobile data is still tried first on home Wi-Fi.
3.3 Network hints
The memory of one device only knows that device's experience. Network hints share the experience of other devices on the same network without identifying anyone.
- Signed network token. The panel gives the app a token containing the country and network number (ASN), signed with HMAC and valid for 48 hours.
- Reporting. The app sends per-protocol results with this token. The request's IP address is never used for observations; while connected, that address is a VPN server anyway.
- Anonymous hourly counters. The panel keeps only hourly counters per (country, ASN, protocol): successes, failures and the number of distinct devices. Distinct devices are counted with per-hour keyed pseudonyms that are deleted when the hour ends. The counts are kept for 7 days.
- Thresholds. A protocol counts as "blocked on this network" when at least 5 devices tried it in the last 24 hours with a success rate below 20 %. If the network has too little data, the country level is used: at least 20 devices and below 10 % success.
- Never all. All offered protocols are never marked blocked.
- Use. The server list response reports the blocked protocols. The app tries them last, unless one worked on this device on this network in the last 24 hours.
3.4 Connect-time health check and the verify path
- A mode only counts as working when real traffic flows. The check asks common connectivity-test addresses (Cloudflare, Google, Microsoft) through the tunnel, never Colitu's own servers. It takes about 4–6 seconds per mode.
- Before a mode is marked as not working, the app checks that the device is online at all, so a dropped Wi-Fi isn't blamed on the mode.
- Time budget: about 8 seconds per mode (start plus traffic check), about 20 seconds per server, at most about 45 seconds for a whole connect. Usually a few seconds are enough.
- Android shows "Connected" only after the check passes; until then it shows "Checking…".
- Verify path. Every time the core starts, a private check input reachable only from the device is added (with a random port and random credentials). The first routing rule sends it straight to the main path. The connect-time check therefore measures the main path only; a broken main path can't hide behind the spare and is avoided on the next connect.
3.5 Server fallback
- In automatic mode, if none of a server's modes pass the health check, the app moves to the next server in the list, up to 3 servers per connect. Meanwhile "Trying another server…" is shown.
- The failed server stays at the back for 30 minutes on that network.
- On Windows and Linux with the kill switch on, the automatic retry moves down the list instead of retrying the same server.
- If every mode fails on a manually chosen server, the error offers a one-tap "Try the fastest server".
3.6 Warm spare
Topology. While connected, the app keeps a second path ready inside the running connection. Traffic passes through a balancer: while the main path is healthy, everything goes over it; if the main path stops working, new traffic moves to the spare by itself. The VPN stays on and no "reconnecting" screen appears. Because the switch happens in the connection itself and not in the app screen, it also works when the app is in the background. The spare only connects when it's used.
Choosing the spare.
| Situation | Spare |
|---|---|
| Automatic mode, a transport is proven on this network (worked in the last 24 hours, on any server) | That transport on the next server in the ranking (same family allowed) |
| Automatic mode, nothing proven | The other family on the next server (UDP ↔ TCP): with a Hysteria2 main path a TCP mode such as VLESS Reality; with a TCP main path Hysteria2 where possible |
| Manual server | Same server, other family; stalled modes are skipped, Shadowsocks last (your location doesn't change) |
| Multihop route or single-mode server | No spare |
Modes and servers that failed on this network are never the spare.
The prefix pitfall. In Xray, balancer and observatory selectors match tags by prefix. If the spare's tag started with the same prefix as the main path's tag, the observatory would treat the spare as the main path too. The spare's tag is therefore chosen so it shares no prefix with the main path.
DNS. Queries from Xray's DNS module have an explicit rule to the balancer.
TCP user timeout. With a spare attached, TCP paths use a 10-second TCP user timeout (the concept from RFC 5482), so dead connections close instead of hanging.
Detection time and cost.
| Measure | Value |
|---|---|
| Detect and switch, Windows/Linux | Usually within about 8 seconds |
| Detect and switch, phones | About 10–13 seconds on average (worst case about 23 seconds) |
| Hysteria2 path on desktop | Immediate when a connection fails; up to about 20 seconds when the path just hangs |
| Extra traffic to check the main path | Phones about 0.3–2 MB per hour, computers about 4–7 MB per hour |
The Windows kill switch allows the spare server too; Linux adds it to its allow-list. The spare's configuration is fetched in parallel with the connect, never delays it, and is cached. Warm spare is on by default; it can be turned off in Advanced mode and is always on in Simple mode.
3.7 Mid-session watcher and the spare-aware rule
- In Adaptive Connect 1.x the mid-session watcher only covered Hysteria2. In 2.0 every transport is watched.
- Cadence: every 5 seconds for the first 90 seconds, then every 30 seconds. Three misses in a row while the device is online mean the main path is considered dead.
- Spare-aware rule. With a spare attached, the watcher checks two things separately: the main path alone, and the normal path (the one the balancer uses).
| Main path | Normal path (incl. spare) | Result |
|---|---|---|
| Working | Working | Nothing happens |
| Dead | Working (spare carries traffic) | No reconnect; the main path is marked for the next connect, and the spare's transport and server lead next time |
| Dead | Dead | Reconnect |
- This rule is the fix for the bug seen in experiment 1 in section 4: the old watcher restarted the connection while the spare was carrying traffic, and broke it.
- Android and iOS reconnect by themselves after an unexpected drop (after about 2, 5 and 15 seconds, then the error is shown) and again when the network comes back. They don't after you press disconnect, sign out, or when the plan or device is paused.
3.8 Stall marks that cannot lock a network
Wrong stall marks could leave a network with no mode to try. Two rules prevent this:
- If the marks would leave at most one of the offered transports, they are ignored for that round.
- Marks found during a round in which everything failed last 10 minutes. The 6-hour lifetime applies only when another transport then carried traffic on the same network, that is, when the problem really was specific to that mode.
3.9 Node online reports and load
Every heartbeat, servers report the number of distinct connected users (Xray's online-user statistics and Hysteria2's online interface). The panel uses complete reports only; if a core's statistics couldn't be read, the previous value is kept. As a result full servers are no longer offered and load-based distribution uses real numbers. These reports are live on all servers.
3.10 Spare health probe and parallel connect
The following two mechanisms are in all apps; both were active in the soak test in 4.3:
- Spare health probe. The spare path is also checked about once a minute through its own check input. A dead spare is only replaced while the tunnel is idle, never during a call.
- Parallel connect ("happy eyeballs"). The best two candidates are checked at the same time, and the first to carry traffic wins. The idea follows the Happy Eyeballs approach of RFC 8305.
4. Evaluation
4.1 Method
| Item | Detail |
|---|---|
| Device | One Android 11 phone |
| Network | One residential Wi-Fi network in a country with active DPI |
| Dates | 9–10 October 2026 |
| Fault injection | At the server, dropping only this client's packets: one protocol (on all listed servers) or one whole server |
| Safety | Rules are tagged and removed by self-deleting timers; nothing else on the server and no other user is touched |
| Traffic probe | Every 5–10 seconds, from the device through the tunnel |
| Output | Outage length per fault |
An internal tool (a fault lab) was written for this: it cuts one protocol (on all listed servers) or one whole server for one client with tagged firewall rules, removes the rules with self-deleting timers, probes the client every 5 seconds for an hour and reports the outage length per fault. In the warm spare experiments in 4.2 faults were applied one at a time; in the soak test in 4.3 the tool injected faults in random order for 60 minutes.
4.2 Results
Ranking. A server in the user's own country had the lowest ping (13–15 ms) but was correctly placed after a nearby foreign server (16–17 ms). In Adaptive Connect 1.x the list came in database-id order, and a server of about 222 ms on another continent was shown as "Recommended".
Freeze observation. On this network every TCP transport to one nearby server (VLESS Reality, VLESS XHTTP, Trojan) froze about 30–50 seconds after connecting: traffic stopped while TCP stayed open. Hysteria2 kept working. The old watcher, which only covered Hysteria2, showed "Connected" with no traffic in this case. The new all-transport watcher detected the freeze and switched transport in about 1.3 seconds.
Baseline. The released version 2.7.0 on the same network ran 100 seconds without interruption on Hysteria2.
Warm spare experiments (the main path's packets were dropped at the server for this client only):
| # | Configuration | Fault | Result |
|---|---|---|---|
| 1 | Same server, spare Trojan | Hysteria2 cut | The spare carried traffic at +12 seconds; then the old watcher restarted the connection and broke it (bug, fixed) |
| 2 | Same server, after the fix | Main path cut | On this network no TCP path survived on that server; recovered about 48 seconds after the block ended (no working path existed during the block) |
| 3 | Automatic mode, spare on the next server chosen as "other family" (Trojan) | Main path cut | Trojan froze too; after a reconnect a Hysteria2 spare carried the traffic |
| 4 | Automatic mode with the proven-transport rule: main Hysteria2 on server A, spare Hysteria2 on server B | Main path cut for 120 seconds | One failed probe at +5 seconds, traffic flowed from +10 seconds to the end; no reconnect, the VPN stayed on. Outage ≈ 5–10 seconds |
4.3 Soak test
On 10 October 2026 a 60-minute soak test was run on the same phone and the same Wi-Fi network, with an Android build containing every 2.0 mechanism (including the spare health probe and parallel connect), in automatic mode.
- Method. The fault lab injected 11 faults on 8 servers in random order, with 2–5 minutes between faults and each fault lasting 60–120 seconds. Two fault types were used: one protocol cut on every server for this client, and one server fully cut for this client.
- Probe. An HTTP request through the tunnel every 5 seconds.
- Result. 741 of 750 probes succeeded (98.8 %).
| # | Fault | Length | Longest outage | Traffic back after the fault ended |
|---|---|---|---|---|
| 1 | VLESS Reality cut on all servers | 83 s | 0 s | 1 s |
| 2 | One server in Germany fully cut | 95 s | 0 s | 3 s |
| 3 | The server in use (primary) fully cut | 101 s | 5 s | 5 s |
| 4 | Shadowsocks cut on all servers | 102 s | 0 s | 4 s |
| 5 | VLESS Reality cut on all servers | 86 s | 0 s | 4 s |
| 6 | One server in Finland fully cut | 65 s | 0 s | 1 s |
| 7 | Shadowsocks cut on all servers | 92 s | 0 s | 2 s |
| 8 | Trojan cut on all servers | 65 s | 0 s | 1 s |
| 9 | One server in Estonia fully cut | 120 s | 0 s | 0 s |
| 10 | Another server in Finland fully cut | 95 s | 0 s | 1 s |
| 11 | Hysteria2 cut on all servers | 61 s | 40 s | 2 s |
Interpretation.
- Fault 3 is the key case: the server in use disappeared for 101 seconds; the user saw a stall of about 5 seconds and the VPN stayed on.
- Fault 11: on this network every TCP protocol is frozen by DPI. With Hysteria2 cut everywhere no working protocol existed; traffic returned 2 seconds after the fault ended.
- Faults on protocols and servers not in use show 0 seconds, as expected: faults elsewhere don't disturb the session.
Limits. One device, one network, one hour; the faults were injected at the server side, not by a real censorship system.
The tested mechanisms are the same in all apps (Android including TV, iOS, Windows and Linux); the measurement was made on one Android phone.
4.4 Discussion
- Experiment 1: A warm spare alone is not enough; the watcher has to know about the spare. This experiment led to the spare-aware rule in 3.7.
- Experiment 2: A spare on the same server is useless when the network freezes all TCP modes of that server (assumption 2). No working path existed here; the measured time is the block itself, not a delay of the mechanism.
- Experiment 3: The "other family" rule is not right on every network. On this network the TCP family froze as a whole; picking the other family meant picking a frozen family. This experiment led to the "proven transport" rule in 3.6: a transport proven to work on this network becomes the spare on the next server.
- Experiment 4: With a proven transport and a different server combined, cutting the main path reached the user only as a short pause.
- Soak test: A 101-second full cut of the server in use came down to a stall of about 5 seconds; the only long outage was fault 11, when no working protocol was left.
- Freeze observation: The gap between "Connected" and "traffic flows" is a real failure type, and watching only one transport is not enough.
4.5 Not yet measured
Switching between Wi-Fi and mobile data and measurements of iOS, Windows and Linux on real devices have not been done yet.
5. Limitations
- Exit rotation. This feature forces VLESS on one server and disables the warm spare. On the measured network this meant repeated freezes. This is a documented limitation.
- Open TCP streams. During a switch to the spare, calls continue after a short glitch, and pages and messages reconnect on their own; but a large download running at that moment may restart. Truly seamless continuation of downloads is future work.
- Single device, single network, one hour. The evaluation used one Android phone on one network, with a test of at most one hour; the faults were injected at the server side, not by a real censorship system. The results can't be generalised to other networks.
- Stale network key. While the VPN is on, the panel can't see which network the request comes from; the app uses the last network it learned until the list is fetched again with the VPN off.
- Online counts. Users connected through sing-box are not included in the online reports; a user connected to two servers is counted twice in the total.
- No spare. There is no warm spare with multihop routes or on single-mode servers. sing-box can't run XHTTP.
6. Future work: the CSL session layer
The Colitu Session Layer (CSL) is a session layer planned on top of the warm spare. The plan is a draft, and no date or result is promised.
- What it adds. Today the warm spare moves new connections to the spare, but open downloads and calls carried over TCP can break. With CSL the app opens one session that survives even when the tunnel underneath (the carrier) changes. In the lab prototype the longest pause was 0.08–1.75 seconds, and downloads were verified to continue unbroken.
- Its limit. A CSL session is bound to one server; the first version can't move a session to another server. CSL solves problems within the same server (switching between Wi-Fi and mobile data, a protocol being blocked, a momentary drop). If the server itself goes down, the warm spare takes over. The two don't replace each other; they work together.
- Phases. First, real-tunnel tests (Reality, Trojan, Hysteria2) on temporary test servers with drop, blocking, network-change and load scenarios; then the server-side service, a per-server on/off flag in the panel and an emergency switch that turns all of CSL off; then Windows first, followed by Linux, Android and iOS last. The targets include a pause of no more than 2 seconds on a carrier change while a spare is ready, and at least 90 % of the speed of a direct connection.
- UDP and calls. The first version protects TCP only; UDP (WhatsApp and Meet calls, QUIC/HTTP3) stays under the warm spare's protection as today. The next step is to carry UDP as datagrams inside the CSL session; a gap shorter than 2 seconds is expected on a carrier change. An optional "force QUIC to TCP" setting in Advanced mode is also being considered.
- Without CSL. If the flag is off on the server, CSL doesn't answer, or it says "busy" or "refused", the client goes back to today's path within a few seconds. A session left without a carrier for more than 90 seconds is lost. A watchdog shuts down a stuck CSL client.
- How users get it. As "Seamless session (beta)" in Advanced mode; off by default and shown only on servers with the flag on.
7. Conclusion
The core lesson of Adaptive Connect 2.0 is that real traffic, not ping, decides whether a path works, and that this decision depends on the network. Ranking that respects proximity and privacy, network memory kept on the device, anonymous network hints and a health check that measures the main path only make it possible to find the right path. While connected, the warm spare together with a watcher that covers every transport and knows about the spare turns a cut main path into a pause of about 5–10 seconds. In the 60-minute soak test 98.8 % of the probes succeeded, and a full cut of the server in use reached the user as a stall of about 5 seconds. The evaluation is narrow; network changes and measurements on the other platforms are the next steps.
References
- RFC 8305, Happy Eyeballs Version 2: Better Connectivity Using Concurrency.
- RFC 9000, QUIC: A UDP-Based Multiplexed and Secure Transport.
- RFC 5482, TCP User Timeout Option.
- Hysteria2 project documentation.
- Xray-core project documentation.
Still need help?
Colitu Bot answers in seconds, and our support team helps in three languages.
