- Assign ETCP_ID_NAT_DETECTION = 0x02, etcp_bind in create, etcp_unbind in destroy
- nat_detection_receive_cbk: own packet dispatcher (PING_REQ/RESP, NAT_INFO/CHECK_REQ)
- 4 handle_* become static internal, send_nat_check_req stays static (dead code)
- Rename TOPO_* -> NATDET_*: all subcommands, structs, message types
- Remove all NAT/PING constants and structs from topo_group.h
- Remove dispatch lines and subcmd_name cases from topo_group_receive_cbk
- API: clear sections (lifecycle/internal service/control/test helpers)
- trigger_checks: marked with IMPORTANT note about call order (must be after add_to_senders)
Root cause: db_handle_init_resp marked sync_state=2 when my_dh==peer_dh
without checking that all my records fit in the confirmed prefix.
If mc > tp+1 (I have records beyond the common prefix), the tail was
never sent — lost records.
Fix:
- Remove special-case 'sc==0 && my_dh==peer_dh' (redundant, same bug)
- In general dh_match: if mc > tp+1, send remaining records [tp+1,mc-1]
using the same inline loop pattern as 'peer empty' branch
Test: 3 phases covering all dh_match paths:
- Phase 1: B empty+late instance → dh_match at tp=0 (sc=0) → tail-send 49
- Phase 2: A=80 B=50 → dh_match at tp=49 (sc>0, sparse) → tail-send 30
- Phase 3: fresh instances → peer empty → send all 30
- db_sync_on_conn_up: per-instance synced/skipped counters
- db_sync_peer_check_cb: summary INFO of check results
- db_sync_instance_add: peers_found + peers_synced counts
- db_handle_init_sync/initiate_sync: enhanced with tbl name
- db_handle_init_resp: tp + my_mc in sync complete/dh match
- db_handle_send_data: range + SYNC_DONE details; stop reason when mc>pk
- db_handle_sync_done: both counts + datahashes in mismatch; dh in confirm
- db_sync_recv_cb: INFO when instance NOT FOUND (timing race)
- cs_handle_channel_info_resp: sync readiness after channel ready
- cs_handle_welcome: note about db_sync peer_check timer
- on_db_sync_insert: counter (#1-3 then every 10) with ch/n/ts/dh/ct
Priority: iterate instance->connections for links with NAT_TYPE_DIRECT
(connected and reachable). Fallback: own socket addresses.
Removed broken getOnlinePeerAddresses() which returned other nodes' addresses.
etcp_conn_set_peer_node_id at state=0 updates conn->peer_node_id but
does not reindex the queue entry (key stays 0). Check the queue entry
directly to detect unindexed outbound connections.
- chat_core.c: CREATE TABLE IF NOT EXISTS peers_* in ensure_channel_ready
- utun_node.cpp: GUI_ERROR redirected to debug log file via guiDebugFile()
- etcp_connections.c: reuse unindexed outbound conn (peer_node_id==0)
on cross-connect INIT to avoid 'Can not insert link to socket'
Root cause: accept queue overflow (backlog=32) caused kernel SYN drops.
Children waited ~1s for SYN retransmit timeout (5 gaps × 1s = 5s out of 7.2s).
With backlog=1024 the timeout drops from ~7.2s to ~0.065s (110× faster).
Router deliver now visible without trace — shows svc_id, len, remote_node_id, conn ptr, callback ptr
This will confirm whether CHANNEL_INFO_REQ reaches chat_sync on the server side
- cs_find_conn_for_node: show queue state, entry found/missing, conn flags
- cs_on_conn_up: log when skipping due to !initialized
- etcp_conn_queue_set_ready: log reindex details (old/new entry, count)
- etcp_conn_set_peer_node_id: show state+reindexed flag
- utun_instance.h: add chat_connections queue (indexed by node_id, 8 bytes)
- chat_sync.c: queue_new in init, queue_free in destroy, queue_find_data_by_index in cs_find_conn_for_node
- chat_sync.c: remove from queue in cs_on_conn_down
- chat_core.c: add to queue in connect_from_invite and connect_auto
- chat_core.c: remove from queue in cc_parallel_cleanup and ca_cleanup
- db_sync.c: replace linear conn search with queue_find_data_by_index (db_sync_send_hash)
- etcp_connections: keepalive SEND/RECV/DECRYPT_OK: DEBUG→TRACE; DECRYPT_FAIL: WARN→DEBUG
- db_sync_on_conn_up: skip sync if !conn->initialized, wait for next UP event
- db_sync peer_check_timer+instance_add: also check conn->initialized
- chat_sync cs_on_conn_up: ignore conn_up before initialized (cs_send fails otherwise)
- db_sync_on_conn_up log now includes init= and links=
- Replace db_sync_stub.c with real src/db_sync.c in libutun
- Remove message sync protocol from chat_sync (INIT_SYNC/INIT_RESP/SEND_DATA/PUSH/ACK_PUSH/SYNC_DONE)
- Keep chat_sync auxiliary: join protocol, peer management, auto-connect
- Remove chat_core_insert_record, cursor_*, mark_sent/ttl_delete stubs
- chat_core on_db_sync_insert now only does SQLite INSERT + GUI notify
- Move db_sync_init() from instance_init_common to utun_instance_init for config override timing
- Force db_sync_enabled=1 in utun_node.cpp
- db_sync handles all P2P message sync via service 0x20
Replace etcp_router_bind/etcp_route_send/etcp_router_unbind with
direct etcp_bind/etcp_send/etcp_unbind + find-conn-by-node_id helpers.
Same approach as db_sync (0x20). Service IDs unchanged (0x30, 0x31).
When etcp_conn_reinit is called, it clears ETCP queues (input_wait_ack
etc.) which frees dgram objects to data_pool. If a router retransmit
timer fires concurrently, it sends through the same conn → normalizer
allocates from corrupted data_pool → SIGSEGV.
Fix: call etcp_router_pause_retrans_for_node() BEFORE etcp_conn_reset()
to cancel retrans timers and free inflight_q entries for all router
connections to the peer. No callbacks, no close notifications —
lightweight pause, router connections stay alive and resume when
ETCP connection comes back up.