ER7412-M2 — reproducible 32-second periodic datapath stall
Subject: ER7412-M2 — reproducible 32-second periodic datapath stall (300 ms to over 3 seconds) affecting all traffic through the device, still present on firmware 1.2.1
Summary
Our ER7412-M2 stops forwarding traffic at a fixed interval of 32.0 seconds. Every device behind it becomes unreachable for the duration of the stall. Measured with a 3000 ms ICMP timeout, the stall is a median of 609 ms, a 90th percentile of 984 ms, and in the worst events it exceeds 3 seconds — a complete outage of the datapath, including the gateway's own LAN interface. This happens with the LAN and WAN essentially idle (1 Mbps LAN, 37 Mbps WAN on a 1 Gbps link) and with gateway CPU reported as 0–4%.
Four controlled experiments narrow this down sharply: the delay affects even pure Layer-2 bridged traffic through the device (no routing involved), it is identical over a completely different physical uplink, it is unchanged immediately after a reboot, and it is unchanged after upgrading to the current firmware 1.2.1 Build 20260811. Combined with twelve eliminated hypotheses, this points to something in the device's common datapath — switch-chip driver, bus servicing or firmware scheduling — rather than to routing, configuration or cabling. We are asking for confirmation and a fix.
Environment
| Item | Value |
|---|---|
| Gateway | ER7412-M2(UN) v1.20, firmware 1.2.1 Build 20260811 Rel.72892 (previously 1.1.0 Build 20251015 Rel.63594 — behaviour identical on both), S/N 2256281000700, MAC 0C-EF-15-AC-C1-5C |
| Controller | OC200, Omada Controller 6.2.14.12 |
| Switches | TL-SX3016F (core), 2× TL-SG3452X, TL-SG2428P, TL-SG2008 |
| APs | 4× EAP653(EU) v1.0, firmware 1.3.11 Build 20260703 Rel. 16605 |
| Networks | VLAN 101 Servers 10.0.1.0/24 · VLAN 102 Clients 10.0.2.0/24 · VLAN 103 Guest 10.0.3.0/24 · VLAN 104 Cameras 10.0.4.0/24 |
| WAN | Dual WAN: 2.5G WAN/LAN1 (primary) + WAN/LAN3 (backup), link backup enabled |
| VPN | L2TP/IPsec server, 8 accounts, pool 10.0.10.0/24 |
| Gateway ↔ core link | 1 Gbps |
All inter-VLAN routing is currently performed by the gateway.
Symptom and measurement method
From a Windows client in VLAN 102 we ping four targets simultaneously (asynchronous ICMP, one round per second, 32-byte payload, 1000 ms timeout) and log every round in which any target exceeds 50 ms:
10.0.1.1— gateway LAN interface10.0.1.70— server in VLAN 101 (crosses the gateway)10.0.1.71— second server in VLAN 101 (crosses the gateway)10.0.2.66— control target in the same VLAN as the client (does not cross the gateway)
Key observations:
- The three targets behind the gateway spike simultaneously and by the same amount, within a few milliseconds of each other.
- The gateway's own address is consistently 20–60 ms faster than the two servers — exactly one extra hop.
- The control target
10.0.2.66never exceeded 50 ms — not once in roughly two hours of measurement across all sessions. - The same spikes are observed simultaneously from two independent client machines.
- Interval between spikes is a strict multiple of ~31.1 s. In one session all 16 intervals were exact multiples with zero exceptions.
This isolates the delay to the gateway itself.
Measurement data
| Session | Duration | Events | Period | Median (gateway) | Max | Notes |
|---|---|---|---|---|---|---|
| 19.08 12:55–13:27 | 32.3 min | 35 | 31.19 s | 368 ms | 852 ms | normal working hours |
| 19.08 22:33–22:52 | 19.1 min | 16 | 30.95 s | 455 ms | 694 ms | office empty, zero active VPN sessions |
| 20.08 08:33–08:53 | 19.5 min | 11 | 30.68 s | 228 ms | 386 ms | OC200 controller powered off |
| 21.08 10:13–10:32 | 18.8 min | 12 | 31.11 s | 449 ms | 793 ms | baseline |
| 21.08 11:15–11:35 | 19.7 min | 17 | 31.10 s | 488 ms | 955 ms | all 13 cameras stopped |
| 21.08 12:22–12:41 | 18.8 min | 34 | 32.27 s | 612 ms | 957 ms | ICMP + TCP measured together |
| 21.08 13:04–13:23 | 19.1 min | 39 | 31.87 s | 582 ms | 916 ms | Layer-2 test, laptop behind the gateway |
| 22.08 08:32–08:50 | 18.3 min | 15 | 32.31 s | 607 ms | 808 ms | copper uplink via another switch |
| 22.08 09:01–09:21 | 20.0 min | 27 | 32.27 s | 606 ms | 799 ms | immediately after reboot |
| 01.09 12:12–12:31 | 19.2 min | 31 | 32.02 s | 660 ms | 837 ms | after upgrade to firmware 1.2.1; in 45% of events all three targets behind the gateway were lost at once (1000 ms timeout) |
| 01.09 12:38–12:58 | 19.7 min | 33 | 31.95 s | 609 ms | >3000 ms | firmware 1.2.1, ICMP timeout raised to 3000 ms to measure the true stall duration |
Detection threshold is 50 ms, so only the larger spikes are counted; the underlying event occurs every 32 s regardless. Sessions up to and including 01.09 12:31 used a 1000 ms ICMP timeout, which truncates the distribution — the final session was repeated with a 3000 ms timeout specifically to measure how long the stall really lasts.
Session of 01.09 12:38–12:58, 3000 ms timeout — true amplitude of the stall:
| Metric | Value |
|---|---|
| Period | 31.95 s (median), range 31.64–32.98 s, 31 of 32 intervals on the grid |
| Median stall | 609 ms |
| 90th percentile | 984 ms |
| Maximum measured | 1231 ms |
| Probes exceeding 3000 ms | 8 of 99 |
| Events with at least one target ≥1000 ms | 7 of 33 (21%) |
| Complete outages (all three targets lost >3000 ms simultaneously) | 12:53:17.958 — every device behind the gateway unreachable for over 3 seconds |
| Control target in the client's own VLAN | 0 events |
| Server minus gateway latency | median +62 ms, range +7…+152 ms — one extra hop |
The gateway's own LAN interface (10.0.1.1) is affected at the same moment and by almost the same amount as the hosts behind it, so the stall is not limited to forwarded traffic — the device's own control plane stops responding too.
Occasionally the phase resets (an interval that is not a multiple of 31.1 s appears), which suggests a self-rescheduling internal loop rather than a wall-clock scheduled task.
Four decisive experiments
1. Layer-2 only traffic is affected identically (routing is not involved). A laptop was connected to a free LAN port of the ER7412-M2 with a static address in VLAN 101 (10.0.1.101) — the same VLAN as the server doing the measurement. Traffic between them is bridged by the gateway, not routed: no NAT, no session table, no firewall, no routing lookup.
From the server we pinged the gateway itself (10.0.1.1) and the laptop (10.0.1.101) simultaneously for 19 minutes:
- gateway: 21 samples, median 582 ms, max 916 ms
- laptop behind the gateway: 23 samples, median 590 ms, max 919 ms
- in the 18 rounds where both answered, the difference was a median of +1 ms, full range −1 to +9 ms (839/840, 886/887, 916/919, 859/860 …)
- control target in the same VLAN but not behind the gateway: zero events
So the stall is not in the routing path. It affects hardware bridging inside the device as well.
2. A completely different physical uplink changes nothing. The gateway was disconnected from the fibre uplink (SFP, core switch port 8) and connected instead by copper RJ45 to a different access switch, which reaches the core over 10 GbE. Fibre, both transceivers and both original ports were removed from the path entirely.
| Fibre uplink | Copper uplink via another switch | |
|---|---|---|
| Period | 32.27 s | 32.31 s |
| Intervals on the grid | 26 of 33 | 14 of 14 |
| ICMP gateway, median | 612 ms | 607 ms |
| TCP to server, median | 264 ms | 280 ms |
Identical to within a millisecond. Optics, transceivers and ports are excluded.
3. A reboot changes nothing. Measurement started immediately after rebooting the gateway: period 32.27 s, ICMP median 606 ms, 25 of 26 intervals on the grid, TCP affected in 20 of 27 events. The behaviour is present from the first minutes of uptime, so it is not a leak or accumulated state.
4. Upgrading to the current firmware changes nothing. On 31.08 at 22:05 the gateway was upgraded online through the controller from 1.1.0 Build 20251015 Rel.63594 to 1.2.1 Build 20260811 Rel.72892 (DEVICE_ONLINE_UPGRADE ... SUCCESSFUL in the site audit log); the controller re-applied the full configuration at 22:12. A 19-minute measurement on the following day:
| 1.1.0 (after reboot) | 1.2.1 | |
|---|---|---|
| Period | 32.27 s | 32.02 s |
| Intervals on the grid | 25 of 26 | 30 of 30 — every interval an exact multiple of 32.0 s |
| ICMP gateway, median | 606 ms | 660 ms |
| Maximum | 799 ms | 837 ms |
| Events with at least one lost probe | 74% | 65% |
| Events where all three targets behind the gateway were lost simultaneously | — | 45% |
| Control target in the client's own VLAN | 0 events | 0 events |
The behaviour is not merely unchanged: in nearly half of the events the stall now exceeds the 1000 ms ICMP timeout, so its true duration can no longer be measured with this timeout. Note that 1.2.1 contains the full 1.2.0 changeset, including the entry "Optimized the resource consumption of traffic statistic", which we had hoped was related to this issue.
What we have ruled out
| Hypothesis | How it was tested | Result |
|---|---|---|
| Attack Defense / Packet Anomaly Defense | All checks disabled except "Block ping from WAN" | No change |
| Load balancing | Disabled | No change |
| WAN online detection / link backup probing | Detection interval set to Disabled entirely | No change |
| Omada controller polling | OC200 powered off for 20 minutes | Period unchanged; amplitude roughly halved (median 449→228 ms, max 793→386 ms) |
| IPsec / L2TP VPN (DPD, rekeying) | Measured during a window with zero active VPN sessions (verified in the site log) | No change |
| User / working-hours load | Measured at 22:33–22:52 with an empty office | No change; median was actually the highest |
| Inter-VLAN camera traffic (13 cameras, ~23 Mbps continuous) | All cameras stopped | No change |
| IntelliRecover | Controller-side feature; already inactive during the powered-off-controller test | No change |
| Routing / NAT / session handling | Layer-2-only path through the device measured (experiment 1) | No change — identical delay |
| Optics, transceivers, uplink ports | Uplink replaced with copper via another switch (experiment 2) | No change |
| Accumulated state / memory leak | Measured immediately after reboot (experiment 3) | No change |
| ICMP rate limiting / control-plane artefact | TCP connect time to port 445 measured in parallel with ICMP | TCP affected too: median 264 ms, max 789 ms, 3 timeouts |
| Firmware version | Upgraded 1.1.0 → 1.2.1 Build 20260811 (experiment 4) | No change; loss became more frequent |
Additional observations
- Nothing appears in the site log at the moment of the spikes. Site log for the measurement windows contains either no entries at all or unrelated IPsec events at different timestamps.
- The CPU graph in the controller shows 0–4%. A sub-second spike every 31 seconds averages to a couple of percent, so this metric cannot resolve the event — we do not consider it evidence either way.
- Switch and disk subsystems on the servers were checked and excluded: intra-VLAN traffic is clean, disk response time on the file server is 0–1 ms, all links are error-free.
- Disabling the controller reducing the amplitude by half, without affecting the period, suggests the periodic task's cost includes controller-related work but the trigger is internal to the gateway.
- Per-client traffic accounting on this installation is numerically broken, and we suspect it is related. Comparing two controller client exports taken exactly four days apart (28.08 06:00 → 01.09 06:00), some counters grow by physically impossible amounts:
| Client | Counter growth over 4 days | Implied sustained rate |
|---|---|---|
| SRV02, download | 3.37 × 10⁸ GB | 7 805 Gbit/s |
| ECO-MD-1C0, upload | 2.90 × 10⁷ GB | 672 Gbit/s |
| SRV02, upload | 2.89 × 10⁶ GB | 67 Gbit/s |
SRV02's stored download counter now reads 4.32 × 10⁹ GB — 4.3 exabytes on a 10 Gbit/s port, two orders of magnitude beyond what the link can physically carry. At the same time a second group of counters is frozen: Cam10, MD027, MD033, MD100 and md004 did not advance by a single byte over the same four days although the devices are online and passing traffic. The statistics subsystem therefore both overflows with garbage and drops updates. This is unchanged on 1.2.1. We raise it because it is objective, reproducible on your side, and concerns the same subsystem named in the 1.2.0 changelog.
Impact
Approximately two user-visible stalls per minute on every flow that crosses the gateway: SMB file access freezes, RDP micro-freezes, VoIP artifacts. Roughly one event in five exceeds one second, and the worst events exceed three seconds, which is long enough to drop TCP sessions and to trigger cluster and monitoring timeouts. With the gateway currently performing all inter-VLAN routing, this affects the entire company continuously.
What we are asking
- Is this a known issue with ER7412-M2 firmware? Which internal process runs on a ~31–32 second cycle and can stall the entire datapath — including hardware bridging — for several hundred milliseconds?
- The issue is present on both 1.1.0 Build 20251015 and the current 1.2.1 Build 20260811. Is a fix planned, and is there an engineering build we can test?
- Is there any setting that disables or reschedules this process?
- Are the corrupted per-client traffic counters described above a known defect, and are the two issues related?
- If diagnosis is needed, we can provide: raw measurement logs from two machines, controller configuration backup, site logs, client exports showing the counter corruption, and a remote session. We can also run a debug build if you can supply one.
We have already scheduled moving inter-VLAN routing off the gateway to an L3 switch as a workaround, but we would like the underlying behaviour understood and fixed, since the gateway will keep handling internet, NAT and VPN.
