When etcp_conn_reinit is called, it clears ETCP queues (input_wait_ack
etc.) which frees dgram objects to data_pool. If a router retransmit
timer fires concurrently, it sends through the same conn → normalizer
allocates from corrupted data_pool → SIGSEGV.
Fix: call etcp_router_pause_retrans_for_node() BEFORE etcp_conn_reset()
to cancel retrans timers and free inflight_q entries for all router
connections to the peer. No callbacks, no close notifications —
lightweight pause, router connections stay alive and resume when
ETCP connection comes back up.
Added callbacks_running flag to ETCP_CONN:
- set to 1 before up_cbks / down_cbks / ready_cbks iteration
- reset to 0 after iteration completes
- etcp_connection_close checks flag and does intentional SIGSEGV
with FATAL log message if called during callback chain
This catches illegal synchronous close from within callbacks
instead of producing cryptic UAF segfaults later.
Protocol: added collision field to INIT_REQUEST_PKT (1 byte before ed25519_pubkey).
When two peers connect simultaneously, the one with is_server=0 (API-created
outbound link) is the master. Resolution by smaller node_id:
- We have master link + smaller node_id → send our INIT with collision=1
- We have master link + larger node_id → yield, process as slave
- Remote sends collision=1 → become slave unconditionally
- session_id change always forces reinit (genuine restart)
Removed unconditional etcp_conn_reinit on INIT_RESPONSE(0x03) — guarded by
reset_done flag. Server-side reinit now has collision check BEFORE reset_done
guard; session_id change penetrates reset_done protection.
Closes cross-connect reinit loop causing endless UP/DOWN flapping.
In etcp_conn_ready / etcp_on_up / etcp_on_down / etcp_connection_create /
tcp_server_on_link, callback chains were iterated with:
while (cbe) { cbe->fn(...); cbe = cbe->next; }
If the callback removes itself from the chain (e.g. ca_ready_cb calls
etcp_conn_remove_ready_cbk which u_free's the entry), cbe->next reads
freed memory → SIGSEGV.
Fixed by saving next pointer before invoking the callback:
while (cbe) { n = cbe->next; cbe->fn(...); cbe = n; }
Problem: when two peers connect simultaneously, each incoming
INIT_REQUEST(0x02) or INIT_RESPONSE(0x03) triggers etcp_conn_reinit
unconditionally, causing endless UP/DOWN flapping loop.
Fix: add reset_done flag to ETCP_CONN:
- 0 at creation and after explicit reinit (reinit allowed)
- set to 1 in etcp_conn_ready (connection stable, block reinit)
Three call sites guarded with !conn->reset_done:
- client: handle_init_response_client (INIT_RESPONSE 0x03)
- server: existing link INIT_REQUEST processing
- server: new link INIT_REQUEST processing
send_reset logic preserved unconditionally — only etcp_conn_reinit
itself is blocked when already stable.
- RECV: log every packet (src addr, link found, session_ready, decrypt result)
- RECV: log init decrypt rejection with packet code (catches 0x03 INIT_RESPONSE
silently dropped by init path)
- SEND: log destination addr+fd in send_init_response and etcp_encrypt_send
- REINIT: log full reason in server and client paths (code, session, got_init,
initialized, links_up state before reinit)
When two peers send INIT simultaneously, the first link occupies the socket.
Incoming INIT arriving after finds the existing outbound link by (ip,port)
instead of failing with 'Can not insert link to socket'.
- chat_core_connect_auto: direct ETCP connect from SQLite (no BGP)
- auto_connect: parallel connect to peers with 3s timeout, retry 10s
- RTT saved to node_addresses on disconnect
- vertical status bar + circle badge with online count per channel
- ChannelListView prevents deselection
- Replace get_time_us() with utun_gettimeofday() for PUSH message timestamps
(db_sync.c, chat_core.c) — monotonic clocks differ across Linux/Windows
- Fix merkle_sync: restart session timer in _handle_request to prevent
double-sided timeout when both nodes start member_sync simultaneously
- Fix INIT_RESP short-format: don't mark synced when my_count > peer_count
- Add db_sync_get_last_timestamp() to avoid duplicate timestamp generation
between db_sync and chat_core on_db_sync_insert
- Fix db_sync_stub to generate timestamps consistently
- build.sh: auto-rebuild chatgui via CMake after utun build
new_conn is only true for the first INIT after server restart.
Subsequent INITs from same peer (recovery loop) see new_conn=0
and revert to old logic, sending 0x05 (no reset) and overwriting
the first 0x03 response.
got_initial_pkt stays 0 until client actually sends seq=1, so
server keeps sending 0x03 across all INIT retries until client
resets and starts from seq=1.
When server restarts, new ETCP_CONN has session_id=0 (calloc).
If client sends INIT_REQUEST_NOINIT (0x04) with session_id=0,
the condition conn->session_id != session_id is false (both 0),
so no reset was sent. Client kept old seq numbers while server
expected seq=1, causing 'Waiting for initial packet' deadlock.
Fix: add || new_conn to the condition — when server created a
fresh connection (was restarted), always send reset to client.
Server now checks both code==ETCP_INIT_REQUEST and session_id mismatch
to decide send_reset. Previously only session_id was checked, so if
session matched (both sides had 00000000), reinit was skipped even
when client explicitly requested reset (code 0x02).
Client side already correct: handle_init_response_client:1473 calls
etcp_conn_reinit if pkt_code==ETCP_INIT_RESPONSE (0x03).
- 1133: DEBUG_DEBUG -> DEBUG_INFO for TX DATA (seq, len, retry, queue lengths)
- 1437: add DEBUG_INFO when initial packet seq=1 accepted
These two INFO lines show the exact order of first sends and receives
to diagnose INIT sequencer sync without guessing.
- 1556: Normal decryption failed -> DEBUG_DEBUG (expected for INIT, no more spam)
- 1581: added 'INIT X25519 OK from %s' with source address
- 1604-1617: INFO log for decrypted packet type (PING/PONG/INIT) with peer_id+src
- 1662-1668: expanded INIT accepted log with all fields:
peer_id, mtu, link_id, socket_id, type, only_local, src_ip:src_port,
session_id, actual src (UDP) — shows NAT mismatch
Without sc_init_ctx, crypto_ctx.initialized==0, sc_set_peer_public_key returns
SC_ERR_NOT_INITIALIZED, and all subsequent sc_encrypt (INIT, etc) fail with
'encryption failed' error 2. Fixed in cm_start_phase_direct and
reverse-connect path (both places that create ETCP_CONN for conn_mgr).
- chat_core_connect_from_invite: add TOPO_SOCKMETA4 (id=0, UNKNOWN, UNKNOWN)
so cm_has_direct_ip/cm_start_phase_direct can match socket_id=0 addr
- conn_mgr: cm_bg_ping_timer_cb reads real addr from nq->node->v4_addrs
instead of NULL; added cm_bg_ping_noop_cb per required etcp_send_ping_to_socket
- invitedialog: set both QClipboard::Clipboard and Selection on Linux/X11
- conn_mgr.c: 5 call sites used queue_entry_new(data_size) with data[] instead
of dgram — all etcp_route_send calls silently dropped. Fixed by switching to
dgram-based allocation (queue_entry_new(0)+u_malloc) matching all other callers.
- DIRCECT_REQ reuses already-allocated pkt as qe->dgram (no extra memcpy).
- etcp_router.c:1078 — 'empty entry' now shows dgram=%p len=%u dst=0x... force=%d
- etcp_connections.c:1139 — 'bad args' now shows all arg pointers to identify NULL
- etcp_connections.c:1145 — 'no socket' now shows addr_family=%d
- conn_mgr.c:633 — 'no candidates' now notes direct+reverse failed
- debug_ui.h: remove NDEBUG guard so GUI_ERROR works in Release builds
- invite_link.h/cpp: add error field to InviteData with specific failure reason
- joindialog: show exact decode error in QLabel instead of generic message
- chat_sync: implement invite connect flow (CHANNEL_INFO_REQ/RESP, JOIN, WELCOME)
- topo_node_sqlite: node/address table init and lookup
- utun_node: getInviteAddresses fallback from ETCP sockets
- minor fixes in mainwindow, chat_core, gui_bridge
- Add CreateGroupDialog with group name input
- Generate X25519+Ed25519 keypairs, random uint64 channel_id,
Ed25519 signature, write to DB via uasync thread
- chat_core_create_channel(): create msg table, topo_node_sqlite_channel_put()
- GUI_EVT_CHANNEL_UPDATED notification to reload channel list
- Per-category debug config: debug_level= + debug_categories= in [gui]
- Unified logging via debug_config: all threads write to same file/console
- debug_enable_console(1) to mirror file output to stderr
- DEBUG_LEVEL_DISABLED=-1: truly disable category, NONE=0 fallback to global
- Fix uasync_post under epoll: process_posted_tasks was never called
- Fix MainWindow constructor parameter ordering (dbPath vs debugFile)
- debug_ui.h: append mode ('a' instead of 'w')
- CHAT groups don't send subnets and skip route insert/delete
- Group type mismatch sends ERR_GROUP_MISMATCH to peer
- New topo_groups_create_group() public API for creating typed groups
- Renamed TOPO_BGP to TOPO_GROUP with topo_group_id (64bit)
- Added TOPO_GROUPS container (ll_queue group_list) in UTUN_INSTANCE
- Memory pools moved from TOPO_GROUP to TOPO_GROUPS (instance-level)
- Default group: TOPO_GROUP_UTUN = 0x8000000000000000
- topo_groups_get_default() searches by group_id, not just first entry
- All topo_bgp_* renamed to topo_group_*; function signatures updated
- Files: topo_bgp.h/c -> topo_group.h/c
- NODEINFO_MSG (protocol, packed): added flags byte + group_id field
- NODEINFO (memory): ref_count, linked list heads instead of inline arrays
- NODEINFO_ROUTES: separate struct with linked list subnet heads
- NODEINFO_Q: node*, routes*, tranzit_data*, hop_list* as separate mallocs
- 6 memory_pools in ROUTE_BGP for NI_* linked list items
- ni_list_count() universal counter via _ni_head cast
- nodeinfo_serialize/deserialize for protocol <-> memory conversion
- NODEINFO_FLAG_SEND_SUBNETS controls subnet data in protocol
- group_id defaults: utun=NODEINFO_GROUP_UTUN(1) with subnets
Fixes:
- deserialization: data+2 instead of data+sizeof(BGP_NODEINFO_PACKET)
- memset: start after ll_entry to preserve ll.size
- hop_src: subtract hop_count*8 to point at correct dynamic offset
- same-ver branch: remove_path(conn) instead of remove_path_by_hop(peer)
- double-free in route_bgp_remove_conn/process_withdraw
- conn_mgr_add_alien_node: rewritten for new structures
- all test files updated for new API
- Add --hardened flag to build.sh with FORTIFY, stack protector, PIE, RELRO
- Add --enable-hardening to configure.ac
- F-001: fix %s -> %.32s to prevent stack leak from non-null-terminated network string
- F-002: add pre-check recv_len >= MAX before buf_space computation to prevent size_t underflow
Added #ifdef __cplusplus / extern "C" { ... } / #endif to ~65 headers
in src/ and lib/, enabling the C library to be linked into C++ code (chatgui).
Fixed packet_dump.h (added missing header guard) and etcp_bbr.h
(#pragma once guard).
- TRANSIT_QUEUE: отдельная очередь на каждую пару src+dst узлов
- ll_queue с хеш-индексом для быстрого поиска transit-очередей
- backpressure через waiter на send_input_q (threshold=0, один пакет за вызов, round-robin)
- Очередь создаётся при первом пакете который не получается отправить напрямую, удаляется при опустошении
- Унифицирован etcp_send: единый путь queue_data_put(send_input_q) для UDP (normalizer->input) и STCP (tx_queue)
- Убран STCP-ветвления из router_forward_transit/drain_cb
- Исправлено перекрытие ll.data и tq->q в TRANSIT_QUEUE (добавлены явные поля src_node_id/dst_node_id)
- Фикс лика dgram в tx_queue_cb/client_tx_queue_cb (добавлен queue_dgram_free)
- transit_queues живут внутри ETCP_CONN, инициализируются лениво, очищаются в etcp_connection_close (до pn_deinit) и stcp_link_close
waiter-based backpressure (commit ba14e31e) could starve HTTPS CONNECT
worker under concurrent 4-worker load — send_q never drained below threshold
fast enough for the waiter to fire before curl timed out (exit=28).
Add retry_timer (500ms, force=1) alongside etcp_router_on_send_ready
waiter: waiter gives instant wake when send_q drains, timer guarantees
delivery via force=1 if waiter hasn't fired in time.
Also improve socks_proxy DEBUG diagnostics in backpressure branches.
- tcp_proxy_server: drain_cb uses etcp_router_on_send_ready instead of retry_timer
- socks_proxy: on_read_cb + tx_waiter_cb use etcp_router_on_send_ready instead of waiter_register
- Result: FIFO round-robin via ll_queue waiters on send_q, all streams get fair share
- No timers — wakeup is event-driven via queue_data_get → check_waiters
- TCP backpressure: read_queue callback_suspended prevents over-reading
process_http_request() теперь обрабатывает любые HTTP-методы:
- CONNECT — туннель как раньше, без изменений
- GET/POST/PUT/HEAD/OPTIONS/... — парсинг абсолютного URL, DNS,
send_connect + пересборка строки с относительным путём + тело DATA
test_socks_http_proxy: +3-й worker (http_proxy через -x без --proxytunnel),
+проверки POST (echo), HEAD (200 OK), OPTIONS (Allow). Все 78 проверок — PASS.