- Add sync_start_tb to SI_PEER, set in db_sync_initiate_sync
- peer_check_cb: after 15s timeout reset sync_state=0 with WARN
- db_handle_error: accept si, reset sync_state=0, DEBUG_ERROR with details
- DB_SYNC_SYNC_TIMEOUT=15 added to db_sync.h
- Fix wrong categories: BGP/DEBUG -> ETCP for connection lifecycle events
- Fix level: WARN -> DEBUG for normal out-of-order pkts before init
- Fix level: INFO -> DEBUG for per-packet TX DATA spam
- Remove duplicate socket init logs (GENERAL duplicates of ETCP)
- Add missing INFO: INIT_RESPONSE received (client), INIT_RESPONSE sent (server)
- Unify all connection establishment messages under ETCP category
- Remove redundant DEBUG address prints in INIT send (info already in INFO msg)
- PING/PONG operations: BGP -> CONNECTION
- DIRECT IP detection: BGP -> NAT
- cs_handle_welcome(): add self to local peers_<ch_id> so
joiner sees own account in channel member list
- cs_on_conn_up(): post CHANNEL_PEERS_ONLINE after reconnect
so GUI shows connected status (was only posted on conn_down)
- etcp_loadbalancer: DEBUG_WARN -> DEBUG_DEBUG for no suitable
link: tmr:wait — normal shaper behavior, not a warning
- chat_core_sync_my_addresses(): remove private IP filter (192.168, 10.x, etc.)
so LAN addresses are stored for direct connections
- topo_node_sqlite_channel_peers_all(): include peers without addresses
in WELCOME message (was skipped entirely, causing empty peer list)
- chat_core_create_channel(): pass real local addresses to member_sync_put()
instead of NULL,0
- chat_core_init(): log number of channels loaded from DB on startup
sc_derive_node_id() hashed the private key, while invite links
(joindialog.cpp) hashed the public key — giving two different
node_ids for the same keypair.
Removed sc_derive_node_id(), kept only sc_derive_node_id_from_pubkey().
All callers updated to use pubkey-based derivation.
- topo_sqlite_db moved from TOPO_GROUPS to UTUN_INSTANCE (direct access)
- topo_groups_set_sqlite_db() removed
- chatgui: chat_sync_connect_node → chat_core_connect_auto (direct ETCP, no conn_mgr)
- chatgui: chat_core_connect_from_invite saves to SQLite → auto connect (no BGP)
- chatgui: removed _on_node_updated BGP callback from member_sync
- chatgui: removed conn_mgr_get_status from status dump
- chatgui: removed all conn_mgr.h and topo_group.h includes
- chatgui: CONN_MGR_ERR_* constants → CC_* in chat_core.h
- removed unused cc_parallel_* code, conn_state_str/conn_type_str helpers
- find_third_node: use nd->inst->connections instead of group->senders_list
- on_conn_up callback auto-calls trigger_checks (no BGP handshake needed)
- Remove group param from link_ready, request_check_all — nat_detection self-sufficient
- Remove trigger_checks + route_ping_cancel_for_conn from public API (internal only)
- topo_group.c: removed last nat_detection_trigger_checks call — completely decoupled
- Header: 4-line summary, no duplicate function list, 7 public functions
- Assign ETCP_ID_NAT_DETECTION = 0x02, etcp_bind in create, etcp_unbind in destroy
- nat_detection_receive_cbk: own packet dispatcher (PING_REQ/RESP, NAT_INFO/CHECK_REQ)
- 4 handle_* become static internal, send_nat_check_req stays static (dead code)
- Rename TOPO_* -> NATDET_*: all subcommands, structs, message types
- Remove all NAT/PING constants and structs from topo_group.h
- Remove dispatch lines and subcmd_name cases from topo_group_receive_cbk
- API: clear sections (lifecycle/internal service/control/test helpers)
- trigger_checks: marked with IMPORTANT note about call order (must be after add_to_senders)
Root cause: db_handle_init_resp marked sync_state=2 when my_dh==peer_dh
without checking that all my records fit in the confirmed prefix.
If mc > tp+1 (I have records beyond the common prefix), the tail was
never sent — lost records.
Fix:
- Remove special-case 'sc==0 && my_dh==peer_dh' (redundant, same bug)
- In general dh_match: if mc > tp+1, send remaining records [tp+1,mc-1]
using the same inline loop pattern as 'peer empty' branch
Test: 3 phases covering all dh_match paths:
- Phase 1: B empty+late instance → dh_match at tp=0 (sc=0) → tail-send 49
- Phase 2: A=80 B=50 → dh_match at tp=49 (sc>0, sparse) → tail-send 30
- Phase 3: fresh instances → peer empty → send all 30
- db_sync_on_conn_up: per-instance synced/skipped counters
- db_sync_peer_check_cb: summary INFO of check results
- db_sync_instance_add: peers_found + peers_synced counts
- db_handle_init_sync/initiate_sync: enhanced with tbl name
- db_handle_init_resp: tp + my_mc in sync complete/dh match
- db_handle_send_data: range + SYNC_DONE details; stop reason when mc>pk
- db_handle_sync_done: both counts + datahashes in mismatch; dh in confirm
- db_sync_recv_cb: INFO when instance NOT FOUND (timing race)
- cs_handle_channel_info_resp: sync readiness after channel ready
- cs_handle_welcome: note about db_sync peer_check timer
- on_db_sync_insert: counter (#1-3 then every 10) with ch/n/ts/dh/ct
etcp_conn_set_peer_node_id at state=0 updates conn->peer_node_id but
does not reindex the queue entry (key stays 0). Check the queue entry
directly to detect unindexed outbound connections.
- chat_core.c: CREATE TABLE IF NOT EXISTS peers_* in ensure_channel_ready
- utun_node.cpp: GUI_ERROR redirected to debug log file via guiDebugFile()
- etcp_connections.c: reuse unindexed outbound conn (peer_node_id==0)
on cross-connect INIT to avoid 'Can not insert link to socket'
Router deliver now visible without trace — shows svc_id, len, remote_node_id, conn ptr, callback ptr
This will confirm whether CHANNEL_INFO_REQ reaches chat_sync on the server side
- cs_find_conn_for_node: show queue state, entry found/missing, conn flags
- cs_on_conn_up: log when skipping due to !initialized
- etcp_conn_queue_set_ready: log reindex details (old/new entry, count)
- etcp_conn_set_peer_node_id: show state+reindexed flag
- utun_instance.h: add chat_connections queue (indexed by node_id, 8 bytes)
- chat_sync.c: queue_new in init, queue_free in destroy, queue_find_data_by_index in cs_find_conn_for_node
- chat_sync.c: remove from queue in cs_on_conn_down
- chat_core.c: add to queue in connect_from_invite and connect_auto
- chat_core.c: remove from queue in cc_parallel_cleanup and ca_cleanup
- db_sync.c: replace linear conn search with queue_find_data_by_index (db_sync_send_hash)
- etcp_connections: keepalive SEND/RECV/DECRYPT_OK: DEBUG→TRACE; DECRYPT_FAIL: WARN→DEBUG
- db_sync_on_conn_up: skip sync if !conn->initialized, wait for next UP event
- db_sync peer_check_timer+instance_add: also check conn->initialized
- chat_sync cs_on_conn_up: ignore conn_up before initialized (cs_send fails otherwise)
- db_sync_on_conn_up log now includes init= and links=
- Replace db_sync_stub.c with real src/db_sync.c in libutun
- Remove message sync protocol from chat_sync (INIT_SYNC/INIT_RESP/SEND_DATA/PUSH/ACK_PUSH/SYNC_DONE)
- Keep chat_sync auxiliary: join protocol, peer management, auto-connect
- Remove chat_core_insert_record, cursor_*, mark_sent/ttl_delete stubs
- chat_core on_db_sync_insert now only does SQLite INSERT + GUI notify
- Move db_sync_init() from instance_init_common to utun_instance_init for config override timing
- Force db_sync_enabled=1 in utun_node.cpp
- db_sync handles all P2P message sync via service 0x20
When etcp_conn_reinit is called, it clears ETCP queues (input_wait_ack
etc.) which frees dgram objects to data_pool. If a router retransmit
timer fires concurrently, it sends through the same conn → normalizer
allocates from corrupted data_pool → SIGSEGV.
Fix: call etcp_router_pause_retrans_for_node() BEFORE etcp_conn_reset()
to cancel retrans timers and free inflight_q entries for all router
connections to the peer. No callbacks, no close notifications —
lightweight pause, router connections stay alive and resume when
ETCP connection comes back up.
Added callbacks_running flag to ETCP_CONN:
- set to 1 before up_cbks / down_cbks / ready_cbks iteration
- reset to 0 after iteration completes
- etcp_connection_close checks flag and does intentional SIGSEGV
with FATAL log message if called during callback chain
This catches illegal synchronous close from within callbacks
instead of producing cryptic UAF segfaults later.
Protocol: added collision field to INIT_REQUEST_PKT (1 byte before ed25519_pubkey).
When two peers connect simultaneously, the one with is_server=0 (API-created
outbound link) is the master. Resolution by smaller node_id:
- We have master link + smaller node_id → send our INIT with collision=1
- We have master link + larger node_id → yield, process as slave
- Remote sends collision=1 → become slave unconditionally
- session_id change always forces reinit (genuine restart)
Removed unconditional etcp_conn_reinit on INIT_RESPONSE(0x03) — guarded by
reset_done flag. Server-side reinit now has collision check BEFORE reset_done
guard; session_id change penetrates reset_done protection.
Closes cross-connect reinit loop causing endless UP/DOWN flapping.
In etcp_conn_ready / etcp_on_up / etcp_on_down / etcp_connection_create /
tcp_server_on_link, callback chains were iterated with:
while (cbe) { cbe->fn(...); cbe = cbe->next; }
If the callback removes itself from the chain (e.g. ca_ready_cb calls
etcp_conn_remove_ready_cbk which u_free's the entry), cbe->next reads
freed memory → SIGSEGV.
Fixed by saving next pointer before invoking the callback:
while (cbe) { n = cbe->next; cbe->fn(...); cbe = n; }
Problem: when two peers connect simultaneously, each incoming
INIT_REQUEST(0x02) or INIT_RESPONSE(0x03) triggers etcp_conn_reinit
unconditionally, causing endless UP/DOWN flapping loop.
Fix: add reset_done flag to ETCP_CONN:
- 0 at creation and after explicit reinit (reinit allowed)
- set to 1 in etcp_conn_ready (connection stable, block reinit)
Three call sites guarded with !conn->reset_done:
- client: handle_init_response_client (INIT_RESPONSE 0x03)
- server: existing link INIT_REQUEST processing
- server: new link INIT_REQUEST processing
send_reset logic preserved unconditionally — only etcp_conn_reinit
itself is blocked when already stable.
- RECV: log every packet (src addr, link found, session_ready, decrypt result)
- RECV: log init decrypt rejection with packet code (catches 0x03 INIT_RESPONSE
silently dropped by init path)
- SEND: log destination addr+fd in send_init_response and etcp_encrypt_send
- REINIT: log full reason in server and client paths (code, session, got_init,
initialized, links_up state before reinit)