Root cause: test1 used lks() (INIT links up) but conn_mgr needs BGP NODEINFO.
Nodeinfo arrives after INIT, so test1 must wait for it.
Also made result volatile to prevent loop optimization.
- Add ed25519_public_key[32] to NODEINFO struct for pubkey distribution
- Add ed25519_public_key[32] to ROUTE_BGP (derived from X25519 privkey at init)
- Add ROUTER_FLAG_SIGNED (0x08) to SVC_ROUTE_HDR flags byte
- router_send_one_flags: when is_signed, sign [hdr][payload] and append 64-byte sig
- etcp_router_recv_cb: verify Ed25519 signature when ROUTER_FLAG_SIGNED set
- New API: etcp_router_conn_send_signed()
- Drop signed packets if sender node not in routing table or no Ed25519 key
- 5 unit tests in test_etcp_router_unit.c covering OK/tampered/unknown/no-key/short
- etcp_conn_reset: queue_resume_callback after clear_queue (input_queue,
input_send_q, input_wait_ack) to prevent suspended callback deadlock
- is_connection_established: check conn->initialized (reliable across
reinit) instead of link->initialized (never cleared on reinit)
- test: force etcp_conn_reinit after server recreation so conn_ok
waits for fresh INIT handshake
- test: guard send_packets() with is_connection_established in reconnect
phases to avoid sending data while conn is down
- etcp_link_update_inflight_lim: self-contained method with clamping,
send_blocked_inflight check, loadbalancer_link_ready on cwnd increase
- etcp_conn_on_inflight_lim_changed: recalc optimal_inflight + resume input
- etcp_ack_recv: use etcp_link_update_inflight_lim instead of manual set
- etcp_request_pkt: fix line 928 check old link (inf_pkt->last_link) not new
- etcp_link_close: recalc optimal_inflight on link removal (both branches)
- etcp_socket_remove: fix dangling pointer loop (nullify after close)
- uasync_print_resources: show timer names, remain time, deleted entries
- diagnostics in utun_instance_destroy and test_etcp_reconnect
External code pushes to write_queue directly (entry_pool + data_pool).
Added queue_set_threshold(write_queue, 32, 0) for write backpressure.
Tests split data into write_chunk_size chunks before queue put.
lib/u_async: uasync_set_socket_read/write — O(1) EPOLL_CTL_MOD toggle per socket
lib/ll_queue: queue_set_empty_callback — one-shot deferred callback on queue drain
lib/tcp_io: TCP connection management via uasync + ll_queue with pool alloc
- read: direct recv into pool buffer, high/low water flow control
- write: autofetch via ll_queue callback, partial send via write_buf
- on_fin/on_error/on_flushed callbacks, FIN deferred until read_queue drained
- getpeername for pre-connected socket detection
tcp_proxy_server: migrated to tcp_io (22→13 fields, 437→270 lines, -7 functions)
tests: test_tcp_io (5 tests), test_tcp_proxy_server rewritten for tcp_conn
- consumer_ack flag: ACK sent only on consumption, not assembly
- etcp_router_consumer_ack() — called from tcp_proxy_client feed_from_transport
- last_ack_sent_tb: interval-based throttle — send immediately if >=10ms passed,
otherwise timer for remaining time
- timer rules: NULL handle after cancel/fire, NULL check before start
- ROUTER_ACK_INTERVAL_TB 1000→100 (100ms→10ms)
- test_etcp_router_unit: updated test 14 for new immediate-send behavior
- tcp_proxy_client_handle_error: set error flag + send CLOSE once if flag was clear
- tcp_proxy_client_handle_data/handle_close: drop silently, don't send ERROR back
- tcp_proxy_server_handle_data/handle_close: drop silently, don't send ERROR back
- Add recv→send 1MB integration test (3 clients × 3 requests, shell + python)
queue_waiter_handle stores defer_q/cb/arg, no extra malloc.
queue_set_waiter_defer(q, 1) enables call_soon dispatch for waiter callbacks.
bbr_integration uses it instead of manual send_waiter_cb wrapper.
rx_acked was used for two conflicting purposes:
1. remote ack of our sends (inflight = tx_seq - rx_acked)
2. our last sent ACK seq (dedup: rx_seq != rx_acked)
When the ack timer fired and set rx_acked = rx_seq, it overwrote
the inflight-tracking value. If rx_seq > tx_seq, the computation
tx_seq - rx_acked underflowed (e.g. 18 - 24 = 0xFFFFFFFA),
permanently blocking router_drain_send_q and causing send_q to
grow indefinitely.
Fix:
- Split rx_acked into tx_acked (remote ack, for inflight) and
last_sent_ack_seq (our ACK, for dedup)
- Incoming ACK handler only advances tx_acked forward (stale guard)
- Use int32_t cast on all inflight comparisons to handle stale states
- Add etcp_router architecture diagram (doc/etcp_router_arch.md)
- New SVC_ROUTE_HDR (packed struct, 22 bytes): cmd+dst+src+seq+svc_id
- ETCP_ROUTER_CONN: state per (remote_node_id, svc_id), hash-indexed in router_conns
- Reorder via recv_q (hash by seq), dedup with 32-bit circular compare
- Periodic ACK (100ms), idle ACK (500ms), inflight limit via tx_seq - rx_acked
- Legacy mode: no conn → direct delivery without reorder
- etcp_router_conn_get/send/close API for seq-managed connections
- Unit test test_etcp_router_unit: 18 tests, 3ms, no ETCP/sockets/BGP
- remove sock_transport, transport abstraction, active connections API
- simplify proxy_conn to 13 fields (was 15), remove tcp_proxy_mapping linked list
- new protocol: stream_id 4B, seq 2B cyclic, HDR_SIZE 8B
- add ERROR subcmd for immediate abort, CLOSE for graceful half-close
- proper half-close: shutdown SHUT_WR, keep reading, no timer
- exhaustive DEBUG_ERROR/WARN on all error/unusual branches
- config parser: report unknown options with filename:line and valid keys
- proxy_recv_cb(p==NULL): tcp_recved(pcb,gap) before closing_tun
- proxy_feed_from_transport: stop when closing_sent
- proxy_poll_cb cleanup: free unsent before tcp_close
- proxy_try_close helper: shared between poll_cb (500ms) and recv_cb
- pending_out check: don't close if data still queued for exit
- Configs: traffic=info, tun=info for diagnostics
- proxy_recv_cb(p=NULL): now calls transport->close() to signal exit side
- Add PROXY FIN/CLOSE/CONNECTED/DATA/FEED/TCP_IN/TCP_OUT logs
- Add RP RECV/SEND/CLOSE logs on exit side
- Fix queue_dgram_free/queue_entry_free ordering (UAF) in all proxy files
- Change TRAFFIC logs from ERROR to INFO level
- Add 2 extra 1MB tests in run_test.sh to verify no post-stress stall
- stress_client.py: add --verbose per-thread timing (connect/send/recv ms)
- mem.c: zero freed memory for early UAF detection
- uasync_set_timeout(0): use FIFO immediate_queue instead of timeout_heap
to guarantee execution AFTER all currently-expired heap timers
- uasync_cancel_timeout: search immediate_queue for timeout=0 timers
- dummynet: start initial shaper with bandwidth-based delay instead of 0
(prevents queue from never filling up when packets arrive sequentially)
- test_dummynet: rename get_time_us -> now_us (conflict with u_async.h)
- test_u_async_comprehensive: fix immediate_timeouts, add timeout_ordering test