memory_pool: +2 байта за объектом — canary 0xAA для детекции buffer overflow,
counter для детекции double-free. При free проверяется canary и counter==0,
при alloc выставляется новый counter.
lwip_tcp: убраны TCP_TW_MAX и счётчики tw_iter (толку нет, краш на 2й итерации).
Оставлены: !pcb->ctx (dangling pointer), self-loop checks, next_owner диагностика.
Добавлена проверка в tcp_alloc: если новый PCB найден в tw_pcbs → halt.
handle_close: tcp_abort→tcp_close + полное обнуление коллбэков (graceful FIN вместо RST)
handle_error: отделён от handle_close, свой cleanup с tcp_abort (RST — нештатная ситуация)
poll_cb/destroy: полное обнуление коллбэков перед tcp_abort (tcp_arg/recv/sent/err/poll)
handle_fin: исправлен лог (shutdown write, не read)
tx_waiter_cb: проверка pc->pcb перед tcp_recved
handle_fin вызывает tcp_shutdown(pcb,0,1) на любом PCB, включая TIME_WAIT.
При active close (FIN_WAIT_1→FIN_WAIT_2→TIME_WAIT) PCB уже в tw_pcbs,
tcp_shutdown на нём портит состояние → tw_pcbs зацикливается → lwip_tcp_input зависает.
Проверка state != TIME_WAIT && != CLOSED перед вызовом.
При CLOSE и FIN от сервера в обратном порядке (FIN relay + CLOSE от двух FIN),
клиент сначала обрабатывает CLOSE (tcp_close → TIME_WAIT), потом FIN
(tcp_shutdown на TIME_WAIT pcb), что портит tw_pcbs список — pcb->next
замыкается на себя, lwip_tcp_input зависает в бесконечном цикле.
tcp_abort шлёт RST и удаляет pcb сразу, без TIME_WAIT.
Shows mtu, s_rand, s_max, offset/xoffset, and whether capped.
RESP: removed old hardcoded 1472+26 cap, replaced by formula.
Both paths use consistent formula based on #defines.
Previous edit accidentally removed tcp_conn_destroy, timer cancellations,
and etcp_router_cancel_send_ready from conn_free. Without tcp_conn_destroy
the socket was never removed from uasync, causing repeated EPOLLERR →
repeated conn_free on freed memory → double-free crash.
Added debug logs for freed guard entry and u_free.
Hardcoded 1472-46 cap exceeded link MTU=1460, causing 'packet too long'
error on INIT handshake. Use link->mtu - UDP_HDR_SIZE - offset - 19
instead of fixed 1472.
error_cb can fire twice for the same fd in one epoll batch,
or conn_free called from multiple callbacks during restart/shutdown.
freed flag makes conn_free idempotent.
tcp_conn_destroy: set read/write queue callbacks to NULL before
queue_free to stop deferred autofetch (drain_cb on freed rc).
write_on_get_cb: throttle consumer_ack to every 8th write_queue
dequeue instead of every single one (was 30+ ACK/sec).
ack_batch field in tcp_proxy_server_conn.
SVC_ROUTE_HDR: +uint8_t flags (23 bytes total).
bit7=START (0x80) — first data packet of new session
bit6=RST (0x40) — sequence violation detected
bit5-4=sess_id (0-3) — cyclically incremented on restart
router_send_one_flags: sets START+sess_id on first data packet.
router_send_ack: sets sess_id only (no START for ACK).
etcp_router_recv_cb: detects restart via sess_id change or START flag.
Old seq=0 heuristic replaced with explicit flags.
etcp_router_conn_get: init sess_id=0, peer_sess_id=0, start_sent=0.
etcp_router_conn_restart: increment sess_id before closing rconn.
etcp_router: new etcp_router_conn_restart(inst, node_id, svc_id) — closes
all rconns for the peer, notifies service via cb(NULL, entry) with
node_id, resets seq state.
seq=0 detection in etcp_router_recv_cb: when rx_seq >= 256 (MAX_INFLIGHT)
and a data packet arrives with seq=0, detect peer restart and reset rconn.
tcp_proxy_client: handle TCP_PROXY_SUBCMD_RESTART — clear all client conns
and server conns for the restarted peer.
Threshold prevents false positives on in-window seq=0 duplicates.
Six key log points at DEBUG level to trace the full ACK chain:
ETCP_SEND — server sends direction A data
ETCP_RECV — received any svc_route packet
SEND_Q_DRAIN — send_q drained after client ACK
SEND_Q_ACKED — server received client ACK (direction A)
CONSUMER_ACK — server sends consumer ACK (direction B)
ACK_SEND — low-level ACK packet sent
router_schedule_ack skips when rx_seq==last_sent_ack_seq (no-op),
so the first consumer_ack call after setting consumer_ack=1 was
silently ignored. Client never got ACK → send_q stuck at 256.
router_ack_do_send always sends the ACK unconditionally.
handle_connect: send first consumer_ack after setting consumer_ack=1
to unblock client router — fixes rq=32/send_q=256 deadlock.
diag_timer_cb: added send_q=N to DIAG line for instant root cause visibility.
read_queue_drain_cb: SOCK:BACKPRESSURE log with send_q count.
write_on_get_cb: SOCK:ACK debug log on each consumer_ack.
1-sec diag_timer_cb prints rq/wq counts+bytes, outstanding pool
allocations, tcp kernel buffer sizes (SO_RCVBUF/SO_SNDBUF).
Started in handle_connect, cancelled in conn_free.
External code pushes to write_queue directly (entry_pool + data_pool).
Added queue_set_threshold(write_queue, 32, 0) for write backpressure.
Tests split data into write_chunk_size chunks before queue put.
lib/u_async: uasync_set_socket_read/write — O(1) EPOLL_CTL_MOD toggle per socket
lib/ll_queue: queue_set_empty_callback — one-shot deferred callback on queue drain
lib/tcp_io: TCP connection management via uasync + ll_queue with pool alloc
- read: direct recv into pool buffer, high/low water flow control
- write: autofetch via ll_queue callback, partial send via write_buf
- on_fin/on_error/on_flushed callbacks, FIN deferred until read_queue drained
- getpeername for pre-connected socket detection
tcp_proxy_server: migrated to tcp_io (22→13 fields, 437→270 lines, -7 functions)
tests: test_tcp_io (5 tests), test_tcp_proxy_server rewritten for tcp_conn
- consumer_ack flag: ACK sent only on consumption, not assembly
- etcp_router_consumer_ack() — called from tcp_proxy_client feed_from_transport
- last_ack_sent_tb: interval-based throttle — send immediately if >=10ms passed,
otherwise timer for remaining time
- timer rules: NULL handle after cancel/fire, NULL check before start
- ROUTER_ACK_INTERVAL_TB 1000→100 (100ms→10ms)
- test_etcp_router_unit: updated test 14 for new immediate-send behavior
- tcp_proxy_client_handle_error: set error flag + send CLOSE once if flag was clear
- tcp_proxy_client_handle_data/handle_close: drop silently, don't send ERROR back
- tcp_proxy_server_handle_data/handle_close: drop silently, don't send ERROR back
- Add recv→send 1MB integration test (3 clients × 3 requests, shell + python)
queue_waiter_handle stores defer_q/cb/arg, no extra malloc.
queue_set_waiter_defer(q, 1) enables call_soon dispatch for waiter callbacks.
bbr_integration uses it instead of manual send_waiter_cb wrapper.
rx_acked was used for two conflicting purposes:
1. remote ack of our sends (inflight = tx_seq - rx_acked)
2. our last sent ACK seq (dedup: rx_seq != rx_acked)
When the ack timer fired and set rx_acked = rx_seq, it overwrote
the inflight-tracking value. If rx_seq > tx_seq, the computation
tx_seq - rx_acked underflowed (e.g. 18 - 24 = 0xFFFFFFFA),
permanently blocking router_drain_send_q and causing send_q to
grow indefinitely.
Fix:
- Split rx_acked into tx_acked (remote ack, for inflight) and
last_sent_ack_seq (our ACK, for dedup)
- Incoming ACK handler only advances tx_acked forward (stale guard)
- Use int32_t cast on all inflight comparisons to handle stale states
- Add etcp_router architecture diagram (doc/etcp_router_arch.md)