- chat_core.c: CREATE TABLE IF NOT EXISTS peers_* in ensure_channel_ready
- utun_node.cpp: GUI_ERROR redirected to debug log file via guiDebugFile()
- etcp_connections.c: reuse unindexed outbound conn (peer_node_id==0)
on cross-connect INIT to avoid 'Can not insert link to socket'
Root cause: accept queue overflow (backlog=32) caused kernel SYN drops.
Children waited ~1s for SYN retransmit timeout (5 gaps × 1s = 5s out of 7.2s).
With backlog=1024 the timeout drops from ~7.2s to ~0.065s (110× faster).
Router deliver now visible without trace — shows svc_id, len, remote_node_id, conn ptr, callback ptr
This will confirm whether CHANNEL_INFO_REQ reaches chat_sync on the server side
- cs_find_conn_for_node: show queue state, entry found/missing, conn flags
- cs_on_conn_up: log when skipping due to !initialized
- etcp_conn_queue_set_ready: log reindex details (old/new entry, count)
- etcp_conn_set_peer_node_id: show state+reindexed flag
- utun_instance.h: add chat_connections queue (indexed by node_id, 8 bytes)
- chat_sync.c: queue_new in init, queue_free in destroy, queue_find_data_by_index in cs_find_conn_for_node
- chat_sync.c: remove from queue in cs_on_conn_down
- chat_core.c: add to queue in connect_from_invite and connect_auto
- chat_core.c: remove from queue in cc_parallel_cleanup and ca_cleanup
- db_sync.c: replace linear conn search with queue_find_data_by_index (db_sync_send_hash)
- etcp_connections: keepalive SEND/RECV/DECRYPT_OK: DEBUG→TRACE; DECRYPT_FAIL: WARN→DEBUG
- db_sync_on_conn_up: skip sync if !conn->initialized, wait for next UP event
- db_sync peer_check_timer+instance_add: also check conn->initialized
- chat_sync cs_on_conn_up: ignore conn_up before initialized (cs_send fails otherwise)
- db_sync_on_conn_up log now includes init= and links=
- Replace db_sync_stub.c with real src/db_sync.c in libutun
- Remove message sync protocol from chat_sync (INIT_SYNC/INIT_RESP/SEND_DATA/PUSH/ACK_PUSH/SYNC_DONE)
- Keep chat_sync auxiliary: join protocol, peer management, auto-connect
- Remove chat_core_insert_record, cursor_*, mark_sent/ttl_delete stubs
- chat_core on_db_sync_insert now only does SQLite INSERT + GUI notify
- Move db_sync_init() from instance_init_common to utun_instance_init for config override timing
- Force db_sync_enabled=1 in utun_node.cpp
- db_sync handles all P2P message sync via service 0x20
Replace etcp_router_bind/etcp_route_send/etcp_router_unbind with
direct etcp_bind/etcp_send/etcp_unbind + find-conn-by-node_id helpers.
Same approach as db_sync (0x20). Service IDs unchanged (0x30, 0x31).
When etcp_conn_reinit is called, it clears ETCP queues (input_wait_ack
etc.) which frees dgram objects to data_pool. If a router retransmit
timer fires concurrently, it sends through the same conn → normalizer
allocates from corrupted data_pool → SIGSEGV.
Fix: call etcp_router_pause_retrans_for_node() BEFORE etcp_conn_reset()
to cancel retrans timers and free inflight_q entries for all router
connections to the peer. No callbacks, no close notifications —
lightweight pause, router connections stay alive and resume when
ETCP connection comes back up.
ca_ready_cb and cc_parallel_ready_cb called etcp_connection_close
synchronously on non-selected connections from within the ready_cbk
callback chain. This corrupts senders_list in topo_group, causing
UAF in topo_group_start_link_nat_check.
Now uses uasync_call_soon (zero-delay timer) to defer the close
to the next event loop iteration, after the callback chain completes
and callbacks_running flag is reset.
Added callbacks_running flag to ETCP_CONN:
- set to 1 before up_cbks / down_cbks / ready_cbks iteration
- reset to 0 after iteration completes
- etcp_connection_close checks flag and does intentional SIGSEGV
with FATAL log message if called during callback chain
This catches illegal synchronous close from within callbacks
instead of producing cryptic UAF segfaults later.
Protocol: added collision field to INIT_REQUEST_PKT (1 byte before ed25519_pubkey).
When two peers connect simultaneously, the one with is_server=0 (API-created
outbound link) is the master. Resolution by smaller node_id:
- We have master link + smaller node_id → send our INIT with collision=1
- We have master link + larger node_id → yield, process as slave
- Remote sends collision=1 → become slave unconditionally
- session_id change always forces reinit (genuine restart)
Removed unconditional etcp_conn_reinit on INIT_RESPONSE(0x03) — guarded by
reset_done flag. Server-side reinit now has collision check BEFORE reset_done
guard; session_id change penetrates reset_done protection.
Closes cross-connect reinit loop causing endless UP/DOWN flapping.
In etcp_conn_ready / etcp_on_up / etcp_on_down / etcp_connection_create /
tcp_server_on_link, callback chains were iterated with:
while (cbe) { cbe->fn(...); cbe = cbe->next; }
If the callback removes itself from the chain (e.g. ca_ready_cb calls
etcp_conn_remove_ready_cbk which u_free's the entry), cbe->next reads
freed memory → SIGSEGV.
Fixed by saving next pointer before invoking the callback:
while (cbe) { n = cbe->next; cbe->fn(...); cbe = n; }
Problem: when two peers connect simultaneously, each incoming
INIT_REQUEST(0x02) or INIT_RESPONSE(0x03) triggers etcp_conn_reinit
unconditionally, causing endless UP/DOWN flapping loop.
Fix: add reset_done flag to ETCP_CONN:
- 0 at creation and after explicit reinit (reinit allowed)
- set to 1 in etcp_conn_ready (connection stable, block reinit)
Three call sites guarded with !conn->reset_done:
- client: handle_init_response_client (INIT_RESPONSE 0x03)
- server: existing link INIT_REQUEST processing
- server: new link INIT_REQUEST processing
send_reset logic preserved unconditionally — only etcp_conn_reinit
itself is blocked when already stable.
- RECV: log every packet (src addr, link found, session_ready, decrypt result)
- RECV: log init decrypt rejection with packet code (catches 0x03 INIT_RESPONSE
silently dropped by init path)
- SEND: log destination addr+fd in send_init_response and etcp_encrypt_send
- REINIT: log full reason in server and client paths (code, session, got_init,
initialized, links_up state before reinit)
- Added struct ca_ctx** ctxs to ca_state (parallel array)
- ca_cleanup now frees all ctxs[i] and the ctxs array
- ca_ready_cb: removed u_free(ctx), added etcp_conn_remove_ready_cbk before cleanup
- ca_timeout_cb: removed u_free(ctx) — ctx lives until ca_cleanup
- All error paths set ctxs[i]=NULL for ca_cleanup safety
etcp_connection_create does NOT add conn to inst->connections list.
Without this, incoming INIT handler can't find the outbound conn
by peer_node_id and creates a new one, causing insert_link collision.
When two peers send INIT simultaneously, the first link occupies the socket.
Incoming INIT arriving after finds the existing outbound link by (ip,port)
instead of failing with 'Can not insert link to socket'.
- ca_timeout_cb: removed etcp_connection_close, ETCP lives past 3s timeout
- ca_state.cancelled flag + chat_core_connect_auto_cancel() for external cancel
- chat_core_connect_auto returns ca_state* via out_state
- Rewritten auto_connect: cursor (ch,peer) instead of bulk collect
- GC every 1s: closes flights with no link_status after 3s
- ac_result_cb no longer clears ca_state (GC handles cleanup)
- stop: cancels all flights via chat_core_connect_auto_cancel
- chat_core_connect_auto: direct ETCP connect from SQLite (no BGP)
- auto_connect: parallel connect to peers with 3s timeout, retry 10s
- RTT saved to node_addresses on disconnect
- vertical status bar + circle badge with online count per channel
- ChannelListView prevents deselection
- Replace get_time_us() with utun_gettimeofday() for PUSH message timestamps
(db_sync.c, chat_core.c) — monotonic clocks differ across Linux/Windows
- Fix merkle_sync: restart session timer in _handle_request to prevent
double-sided timeout when both nodes start member_sync simultaneously
- Fix INIT_RESP short-format: don't mark synced when my_count > peer_count
- Add db_sync_get_last_timestamp() to avoid duplicate timestamp generation
between db_sync and chat_core on_db_sync_insert
- Fix db_sync_stub to generate timestamps consistently
- build.sh: auto-rebuild chatgui via CMake after utun build
chat_sync_push: add chain_hash parameter, use it instead of zeros
(fixes chain_hash mismatch → all PUSH messages were rejected)
on_db_sync_insert: pass computed chain_h to chat_sync_push
cs_handle_channel_join: call member_sync_start for the joiner
(server now initiates member sync when client joins.
Previously merkle_sync was only started by client → timeout
because server never responded)
cs_handle_welcome: add DEBUG_INFO before member_sync_start on client side