- Add sync_start_tb to SI_PEER, set in db_sync_initiate_sync
- peer_check_cb: after 15s timeout reset sync_state=0 with WARN
- db_handle_error: accept si, reset sync_state=0, DEBUG_ERROR with details
- DB_SYNC_SYNC_TIMEOUT=15 added to db_sync.h
Root cause: db_handle_init_resp marked sync_state=2 when my_dh==peer_dh
without checking that all my records fit in the confirmed prefix.
If mc > tp+1 (I have records beyond the common prefix), the tail was
never sent — lost records.
Fix:
- Remove special-case 'sc==0 && my_dh==peer_dh' (redundant, same bug)
- In general dh_match: if mc > tp+1, send remaining records [tp+1,mc-1]
using the same inline loop pattern as 'peer empty' branch
Test: 3 phases covering all dh_match paths:
- Phase 1: B empty+late instance → dh_match at tp=0 (sc=0) → tail-send 49
- Phase 2: A=80 B=50 → dh_match at tp=49 (sc>0, sparse) → tail-send 30
- Phase 3: fresh instances → peer empty → send all 30
Root cause: accept queue overflow (backlog=32) caused kernel SYN drops.
Children waited ~1s for SYN retransmit timeout (5 gaps × 1s = 5s out of 7.2s).
With backlog=1024 the timeout drops from ~7.2s to ~0.065s (110× faster).
Remove secondary level filtering in debug_output(). Now the only filter
is debug_should_output() (per-category + global fallback). Output goes
to file (if file_output set) or console (if console_enabled), with no
additional level check.
Removed:
- debug_set_console_level(), debug_set_file_level()
- console_level, file_level fields from debug_config_t
- debug_set_level() no longer sets console_level
- Renamed TOPO_BGP to TOPO_GROUP with topo_group_id (64bit)
- Added TOPO_GROUPS container (ll_queue group_list) in UTUN_INSTANCE
- Memory pools moved from TOPO_GROUP to TOPO_GROUPS (instance-level)
- Default group: TOPO_GROUP_UTUN = 0x8000000000000000
- topo_groups_get_default() searches by group_id, not just first entry
- All topo_bgp_* renamed to topo_group_*; function signatures updated
- Files: topo_bgp.h/c -> topo_group.h/c
- NODEINFO_MSG (protocol, packed): added flags byte + group_id field
- NODEINFO (memory): ref_count, linked list heads instead of inline arrays
- NODEINFO_ROUTES: separate struct with linked list subnet heads
- NODEINFO_Q: node*, routes*, tranzit_data*, hop_list* as separate mallocs
- 6 memory_pools in ROUTE_BGP for NI_* linked list items
- ni_list_count() universal counter via _ni_head cast
- nodeinfo_serialize/deserialize for protocol <-> memory conversion
- NODEINFO_FLAG_SEND_SUBNETS controls subnet data in protocol
- group_id defaults: utun=NODEINFO_GROUP_UTUN(1) with subnets
Fixes:
- deserialization: data+2 instead of data+sizeof(BGP_NODEINFO_PACKET)
- memset: start after ll_entry to preserve ll.size
- hop_src: subtract hop_count*8 to point at correct dynamic offset
- same-ver branch: remove_path(conn) instead of remove_path_by_hop(peer)
- double-free in route_bgp_remove_conn/process_withdraw
- conn_mgr_add_alien_node: rewritten for new structures
- all test files updated for new API
Root cause: 80kbps congestion causes 9+ second backpressure during push
phase. No packets reach server → keepalive timeout (2s default) fires
→ ETCP link down → BGP cleanup → etcp_router_conn_restart()
- frees send_q (up to 64 = ROUTER_MAX_SEND_Q_PACKETS lost)
- sends 9-byte restart notification to service handler (conn=NULL)
- resets seq to 0
Fixes:
1. Set keepalive_timeout=60000, keepalive_adaptive=0 in test configs
to prevent spurious timeouts during intentional congestion testing
2. Handle 9-byte restart notification in srv_handler:
when conn==NULL && len==9 → reset g_expected_seq=0 gracefully
- tcp_proxy_server: drain_cb uses etcp_router_on_send_ready instead of retry_timer
- socks_proxy: on_read_cb + tx_waiter_cb use etcp_router_on_send_ready instead of waiter_register
- Result: FIFO round-robin via ll_queue waiters on send_q, all streams get fair share
- No timers — wakeup is event-driven via queue_data_get → check_waiters
- TCP backpressure: read_queue callback_suspended prevents over-reading
process_http_request() теперь обрабатывает любые HTTP-методы:
- CONNECT — туннель как раньше, без изменений
- GET/POST/PUT/HEAD/OPTIONS/... — парсинг абсолютного URL, DNS,
send_connect + пересборка строки с относительным путём + тело DATA
test_socks_http_proxy: +3-й worker (http_proxy через -x без --proxytunnel),
+проверки POST (echo), HEAD (200 OK), OPTIONS (Allow). Все 78 проверок — PASS.
uasync: replace raw socket_node* handle with packed index — fixes UAF after socket_array realloc
tcp_io: defer u_free in tcp_conn_destroy via uasync_call_soon — prevents callback-chain UAF
tcp_io: guard read_cb/write_cb with NULL queue checks after deferred destroy
tcp_io: save read_queue to local before queue_data_put + NULL guard after
route_connectivity: linked-list probe_ctx — cancel all parallel probes before nq free
- etcp_connections: fix link leak & UAF in etcp_socket_remove — remove_link inside etcp_link_close shifts array, skip NULL set and i++
- stcp_server: fix UAF in stcp_conn_process_recv — defer free via uasync_call_soon after stcp_conn_do_close
- etcp: save pkt_len before queue_data_put to output_queue (callback may free entry synchronously)
- socks_proxy: add pool bounds check before memcpy, UINT16_MAX truncation guard, freed flag
- tcp_proxy_server: reorder cleanup (cancel waiters before tcp_conn_destroy), UINT16_MAX guard
- test_route6_lib: increase nodes[] array to STRESS_NODES + STRESS_OPS
В одном epoll_wait batch три события для fd: EPOLLERR (освобождает tc),
EPOLLIN (read_cb), EPOLLOUT (write_cb). read_cb/write_cb guard проверяет
tc->sock==SOCKET_INVALID, но sock обнулялся ПОСЛЕ on_error.
Перенос sock=SOCKET_INVALID до on_error защищает от use-after-free.
Доказано: без фикса segfault на tcp_io.c:374 (write_cb → tc->connected).
С фиксом тест работает без крашей.
- pkt_normalizer: save recvpart->len before queue_data_put (may free synchronously)
- ll_queue: save uasync_call_soon handle in queue_waiter_handle, cancel on free
- uasync: generation counter in epoll data.u64 to skip stale events after fd reuse
- uasync: keep fd_to_index mapping after socket_array_remove for inactive nodes
- test: test_uasync_socket_race — 4 children × 200 iterations, currently flaky
- config_parser: tun_enabled в [global] (по умолчанию yes), не дефолтить
tun_ifname если TUN не нужен
- utun_instance: уважать tun_enabled, чистые дефолты tcp_proxy TUN
- socks_proxy: фикс дэдлока — reply сразу после CONNECT, без ожидания DATA
- utun.conf.sample: документирован tun_enabled, tun_ifname
- test_socks_http_proxy: 50 запросов (SOCKS+HTTP), 2 воркера параллельно,
файлы 1k..100k, сверка через cmp, python http.server + curl
Добавлен модуль socks_proxy.c/h — SOCKS5 и HTTP CONNECT proxy на клиенте.
Использует tcp_io для локальных TCP соединений, туннелит трафик через
существующий ETCP-протокол (ETCP_ID_TCP_PROXY/ETCP_ID_TCP_PROXY_CLIENT).
Конфиг [tcp_proxy_client]:
- socks_enabled/socks_addr — включить SOCKS5 сервер на указанном IP:PORT
- http_proxy_enabled/http_proxy_addr — включить HTTP CONNECT сервер
- Можно включать TUN, SOCKS и HTTP прокси одновременно или раздельно
Весь трефик идёт через exit node (tcp_proxy_server), менять exit не пришлось.
Root cause: when bgp->local_node was reallocated (changed=1), old route
entries remained with stale v_node_info pointing to freed memory. The freed
memory got reused by a BGP peer's NODEINFO_Q, causing route_lookup to return
wrong node_id (b3b2b1b0afaeadac instead of local), which then failed at
etcp_route_send with 'no BGP route' and dropped packets to 10.23.2.1.
Changes:
- route_node.c: route_delete(old) before u_free + route_insert(new) after alloc
- utun_instance.c: remove duplicate route_insert (now atomic inside function)
- route_bgp.c: route_delete before freeing old peer NODEINFO_Q
- route_lib.c: routes_overlap now rejects only exact prefix+network duplicates
- tests/Makefile.am: add route_lib.o to test_route6_lib
- keepalive_adaptive=0 disables adaptive keepalive period growth,
keeping it fixed at 200ms for faster timeout detection in tests
- test_bgp_triangle: 5-node triangle topology test (6 phases):
validates alternative path recording (fix#1), cascading node
removal on full isolation, and full recovery
- config generation via utun_instance_create_from_str() with
pre-generated keys — no temp files needed
Root cause: test1 used lks() (INIT links up) but conn_mgr needs BGP NODEINFO.
Nodeinfo arrives after INIT, so test1 must wait for it.
Also made result volatile to prevent loop optimization.
- Add ed25519_public_key[32] to NODEINFO struct for pubkey distribution
- Add ed25519_public_key[32] to ROUTE_BGP (derived from X25519 privkey at init)
- Add ROUTER_FLAG_SIGNED (0x08) to SVC_ROUTE_HDR flags byte
- router_send_one_flags: when is_signed, sign [hdr][payload] and append 64-byte sig
- etcp_router_recv_cb: verify Ed25519 signature when ROUTER_FLAG_SIGNED set
- New API: etcp_router_conn_send_signed()
- Drop signed packets if sender node not in routing table or no Ed25519 key
- 5 unit tests in test_etcp_router_unit.c covering OK/tampered/unknown/no-key/short