| Message ID | 20260202132309.567382-1-ralf@mandelbit.com |
|---|---|
| State | Awaiting Upstream |
| Headers |
Return-Path: <openvpn-devel-bounces@lists.sourceforge.net>
Delivered-To: patchwork@openvpn.net
Received: by 2002:a05:7000:6911:b0:80a:3855:ce6a with SMTP id
o17csp1762671map;
Mon, 2 Feb 2026 05:23:52 -0800 (PST)
X-Forwarded-Encrypted: i=2;
AJvYcCUickX03sHx9fBaWGV2Y+p8EKjqkXNqfnYFL83ifzt1qN6YvdYcpFgYJX1daI8zJ+U2BnhPrY8qcaY=@openvpn.net
X-Received: by 2002:a05:6830:6d0f:b0:7c7:8113:6f6e with SMTP id
46e09a7af769-7d1a536db2amr6821516a34.27.1770038631919;
Mon, 02 Feb 2026 05:23:51 -0800 (PST)
ARC-Seal: i=1; a=rsa-sha256; t=1770038631; cv=none;
d=google.com; s=arc-20240605;
b=WcdONNE/9LSSpF1QwpmW18v7nA543G9B1AdAVGIVxdPTUBCn/7CHQsW7eYyYQDFFMs
yRuu3rl75vCxLWHmr4M/4dfr+PmdcaeWfFyYHnbr0gMLsBPrMCdaXb51sV4X9lv93vCT
Ta9NBOFDzpr/7azIY19U8+8X7Ul8QXOqdf0AqnM+pois5gw38DFUTQ+/xxe50GtRr3ri
jJ+Ua2ZToasijiY5h35l0f96o747F8hyRsaWNVJ2ap+xYcGdbz0nMvKT3EVY08llRgo/
jf5BirIweGewf47t+swq26bll9HyHBp3qNMbL+RrOrlASTeE7HOfq8TEF6jky7DPF0i7
ho8w==
ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com;
s=arc-20240605;
h=errors-to:content-transfer-encoding:cc:list-subscribe:list-help
:list-post:list-archive:list-unsubscribe:list-id:precedence:subject
:mime-version:message-id:date:to:from:dkim-signature:dkim-signature
:dkim-signature:dkim-signature;
bh=TQngdK2YY4ltpOJBvAzRd6b4BkQpAL3a3rY4ygKIT+M=;
fh=bDmbXayvKcQuWZaaz4JM7kgnS3MJBk3QUq2ehqNuBVc=;
b=GULt0yE3NZQXs90pczcXkNJxyOX9OAdTEf2HL2CSo+4OAXe17870vmhpMJeXrETLQe
TmoVfTGLdc18CGMMRtCpj1bGzjDXQt8/9+8YMu4BSg9V2sPkhggthHHrGfQNJd+LYmzK
a+tdSmYGknQFNMMZRNaaFXnTWgaPBsz4rLonFVHpzh/fHR3nhHl/FSTovgHk75fkvGm2
Vr8q8XLxL+AtowGcFYl4Cy7kck0jq/gWmNXRc9He1bCZOh656kS0aYlNt1MekY2CyI2M
B6HlngqMxLniUp27Z0ZSiVVakQN5+TI4KXcgEKPCauqo1FWCpmiOko1otMEVlYPTw4OE
8XAw==;
dara=google.com
ARC-Authentication-Results: i=1; mx.google.com;
dkim=pass header.i=@lists.sourceforge.net header.s=beta
header.b="hK8t/Ais";
dkim=neutral (body hash did not verify) header.i=@sourceforge.net
header.s=x header.b=Ji1DkUHd;
dkim=neutral (body hash did not verify) header.i=@sf.net header.s=x
header.b=I0oAxzaH;
dkim=neutral (body hash did not verify) header.i=@mandelbit.com
header.s=google header.b=CSzkwdOb;
spf=pass (google.com: domain of
openvpn-devel-bounces@lists.sourceforge.net designates 216.105.38.7 as
permitted sender) smtp.mailfrom=openvpn-devel-bounces@lists.sourceforge.net;
dara=neutral header.i=@openvpn.net
Received: from lists.sourceforge.net (lists.sourceforge.net. [216.105.38.7])
by mx.google.com with ESMTPS id
46e09a7af769-7d18c7f8768si9136505a34.167.2026.02.02.05.23.51
(version=TLS1_2 cipher=ECDHE-ECDSA-AES128-GCM-SHA256 bits=128/128);
Mon, 02 Feb 2026 05:23:51 -0800 (PST)
Received-SPF: pass (google.com: domain of
openvpn-devel-bounces@lists.sourceforge.net designates 216.105.38.7 as
permitted sender) client-ip=216.105.38.7;
Authentication-Results: mx.google.com;
dkim=pass header.i=@lists.sourceforge.net header.s=beta
header.b="hK8t/Ais";
dkim=neutral (body hash did not verify) header.i=@sourceforge.net
header.s=x header.b=Ji1DkUHd;
dkim=neutral (body hash did not verify) header.i=@sf.net header.s=x
header.b=I0oAxzaH;
dkim=neutral (body hash did not verify) header.i=@mandelbit.com
header.s=google header.b=CSzkwdOb;
spf=pass (google.com: domain of
openvpn-devel-bounces@lists.sourceforge.net designates 216.105.38.7 as
permitted sender) smtp.mailfrom=openvpn-devel-bounces@lists.sourceforge.net;
dara=neutral header.i=@openvpn.net
DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed;
d=lists.sourceforge.net; s=beta; h=Content-Transfer-Encoding:Content-Type:Cc:
List-Subscribe:List-Help:List-Post:List-Archive:List-Unsubscribe:List-Id:
Subject:MIME-Version:Message-ID:Date:To:From:Sender:Reply-To:Content-ID:
Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc
:Resent-Message-ID:In-Reply-To:References:List-Owner;
bh=TQngdK2YY4ltpOJBvAzRd6b4BkQpAL3a3rY4ygKIT+M=; b=hK8t/AisDzAIZYytNvVmYh6qg3
ssfoDBr5Jnwl32eBLBAZSwtHDv/dfdTTikpY5L7xneHzHzJqFVElX4m70HrMR4iD8mLF9mJZefo2W
FQurmfvHgVttQ7U25X+jerp915sCPlLIjMVoqdh8mjHmeVdgKC2I58BtfAqWCjLt06DQ=;
Received: from [127.0.0.1] (helo=sfs-ml-3.v29.lw.sourceforge.com)
by sfs-ml-3.v29.lw.sourceforge.com with esmtp (Exim 4.95)
(envelope-from <openvpn-devel-bounces@lists.sourceforge.net>)
id 1vmtu2-0008GI-0q;
Mon, 02 Feb 2026 13:23:42 +0000
Received: from [172.30.29.66] (helo=mx.sourceforge.net)
by sfs-ml-3.v29.lw.sourceforge.com with esmtps (TLS1.2) tls
TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.95)
(envelope-from <ralf@mandelbit.com>) id 1vmtu0-0008GA-Kl
for openvpn-devel@lists.sourceforge.net;
Mon, 02 Feb 2026 13:23:40 +0000
DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed;
d=sourceforge.net; s=x; h=Content-Transfer-Encoding:MIME-Version:Message-ID:
Date:Subject:Cc:To:From:Sender:Reply-To:Content-Type:Content-ID:
Content-Description:Resent-Date:Resent-From:Resent-Sender:Resent-To:Resent-Cc
:Resent-Message-ID:In-Reply-To:References:List-Id:List-Help:List-Unsubscribe:
List-Subscribe:List-Post:List-Owner:List-Archive;
bh=9rhiqiHnsEP0TYwk4vvoiu09iAhA+1BkX6FPLuuez0g=; b=Ji1DkUHdVSuPFsN8xpefXkOmp7
DNo8e4ICDTuFdbdtlp5A3s9uDhf3pwwZBcY9lR50eDQ+OIKNlBR1+4dmngZ1wKCnOBQ7/sGmV9/v1
SKimuKX0/yvkuJLCF1lt1VlwcjnsfL5Xd5KKZScnuEblN5qPirHkKEW2G0PA1zYs9c6E=;
DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=sf.net; s=x
;
h=Content-Transfer-Encoding:MIME-Version:Message-ID:Date:Subject:Cc:To:From
:Sender:Reply-To:Content-Type:Content-ID:Content-Description:Resent-Date:
Resent-From:Resent-Sender:Resent-To:Resent-Cc:Resent-Message-ID:In-Reply-To:
References:List-Id:List-Help:List-Unsubscribe:List-Subscribe:List-Post:
List-Owner:List-Archive; bh=9rhiqiHnsEP0TYwk4vvoiu09iAhA+1BkX6FPLuuez0g=; b=I
0oAxzaHQ2HWZNwMH8G1qGlbaXHO3eqq+EyLABWINr4Th18mmSOc5xRwjBSqXmB/oMN4w3QkyWiLRU
ysmgNgByIP1ngcMlPshCnP2srW75oQqUKJK/zM0vxuj0AIGNgcANaL4tpIq9ekrurGobAunsMvNJG
XEhlK7IzYKQ64EbQ=;
Received: from mail-wm1-f49.google.com ([209.85.128.49])
by sfi-mx-2.v28.lw.sourceforge.com with esmtps
(TLS1.2:ECDHE-RSA-AES128-GCM-SHA256:128) (Exim 4.95)
id 1vmtu0-0002Ab-6a for openvpn-devel@lists.sourceforge.net;
Mon, 02 Feb 2026 13:23:40 +0000
Received: by mail-wm1-f49.google.com with SMTP id
5b1f17b1804b1-4801d7c72a5so35228985e9.0
for <openvpn-devel@lists.sourceforge.net>;
Mon, 02 Feb 2026 05:23:40 -0800 (PST)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
d=mandelbit.com; s=google; t=1770038608; x=1770643408;
darn=lists.sourceforge.net;
h=content-transfer-encoding:mime-version:message-id:date:subject:cc
:to:from:from:to:cc:subject:date:message-id:reply-to;
bh=9rhiqiHnsEP0TYwk4vvoiu09iAhA+1BkX6FPLuuez0g=;
b=CSzkwdOba9JRBAFbTMo7ha17CGu3MWGZabHDkXskVIKsSDoo45IpZ23+6Xe5g/y9l5
ezsX0aAQX0OtO2tR7bWzUxECGtkwZLz+v/tS8YCBMksn2uW4iFXJN5pdUI1HGRaifn+l
0dWqci1IWVQAjM8gMBCB+qDht68DgWEf5mJa4hWbagg/iW9WOwCszM8E+WHir8C+wXz6
omEC8yCQY0FuP84qh5WJtVPyS53AJS617ydU91qg/5WqpkenvnBREnxQfPHfC6IqAVhz
rukyhVtq1eMAC22IdftRRDJ5jr+xGAzk9sGwvj/9w6+3F2VPcVSWSb7A+HX4HODPeHN+
pwVg==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
d=1e100.net; s=20230601; t=1770038608; x=1770643408;
h=content-transfer-encoding:mime-version:message-id:date:subject:cc
:to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date
:message-id:reply-to;
bh=9rhiqiHnsEP0TYwk4vvoiu09iAhA+1BkX6FPLuuez0g=;
b=Sutt4QAwS4iXloVQy54/34rrkQ0Juhyg9fqF6XKE1YJOWrVxdQdpG8CT2Lr/kLXA9r
+zfvBy/9G9LVD8yex0MVL950078rCuQI+oKHsxa/T9kjKzit+sxqMlzrUjO1dTBhb7fl
5alGbQsDdFCd3fp+t9kUrKv8nxBo4i3M9eM5inWY35GfQB8N9jQxHyN2dXUF2+Kmac7e
9fTd4/BGXhUgbWkHyExkutST9A5CN+VHCW4HK3eTRFa2KPOyydaLy1bK3/pQM4csbXpC
IrQlo1OfvrrAJuwc7SXkHkxMgmXakz//tARGEVQ/rjOUatzPP7k2eI1A+DtfHjxRbzZu
Kb3w==
X-Gm-Message-State: AOJu0YzFDeRM3DbRtZx7IjBuAW35ncFyFZaz/Ip1HIawrc4uidcxnN6N
yvdcNwdc/ld+3K0EmhuVQmSHo1cve+jHJJV1CABSFVyF80HzaSlwZwznD7fG999ZMt8/PuciQm5
5cuCX
X-Gm-Gg: AZuq6aIPDZ2VbYPOJzWinDmg+SXk3mRYb0fxmg4nbBaEqoMpbDh+cfLkGkKG8iRMd7R
yBFuJSkq7A4jnN/yi8zHpsMaT68h55XMP/opXQWiDVcLqN0cn/dJYeMblka+I8W3orO9Tq3fhqN
oeINOOniKtx+oiCRcjl4HBlo5enXje9gB74NdZJ2gArTCTHq7jyBOWqsx7CsRROTu9l1iZTgPHP
RECkOqQlLmwR5W1t7RqDpz9YPVI0TQ/u3Bd8SafUORYYf9V8DWIGY/lkXZH2KtsxqulffbF9Xfv
JSaulICOsdZHgYvbFut7TE+pvMClzRXmPPLcYaORs0CwqYnlc23B6Qd4/+4JVIEhCqPFBNSCusg
dSliZDqvfCX0cP1h83SrFAu0/4h93MMvV90rvhGDXa9KyeZ/PzSwmQ86O4e9vxzgl32Nkn7AAM9
/E6nyCQw==
X-Received: by 2002:a05:600c:608e:b0:45c:4470:271c with SMTP id
5b1f17b1804b1-482db4d8210mr129311375e9.18.1770038608148;
Mon, 02 Feb 2026 05:23:28 -0800 (PST)
Received: from fedora ([2a01:e11:600c:d1a0:3dc8:57d2:efb7:51a8])
by smtp.gmail.com with ESMTPSA id
5b1f17b1804b1-48066c37420sm494626275e9.9.2026.02.02.05.23.27
(version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256);
Mon, 02 Feb 2026 05:23:27 -0800 (PST)
From: Ralf Lici <ralf@mandelbit.com>
To: openvpn-devel@lists.sourceforge.net
Date: Mon, 2 Feb 2026 14:23:09 +0100
Message-ID: <20260202132309.567382-1-ralf@mandelbit.com>
X-Mailer: git-send-email 2.52.0
MIME-Version: 1.0
X-Spam-Score: -0.2 (/)
X-Spam-Report: Spam detection software,
running on the system "sfi-spamd-1.hosts.colo.sdot.me",
has NOT identified this incoming email as spam. The original
message has been attached to this so you can view it or label
similar future email. If you have any questions, see
the administrator of that system for details.
Content preview: Currently,
when a connected peer expires (no packets received
within the keepalive interval), we remove the peer and notify userspace of
the deletion. We then insert the peer in the release list and p [...]
Content analysis details: (-0.2 points, 5.0 required)
pts rule name description
---- ----------------------
--------------------------------------------------
-0.1 DKIM_VALID Message has at least one valid DKIM or DK signature
0.1 DKIM_SIGNED Message has a DKIM or DK signature,
not necessarily valid
-0.1 DKIM_VALID_AU Message has a valid DKIM or DK signature from author's
domain
-0.1 DKIM_VALID_EF Message has a valid DKIM or DK signature from
envelope-from domain
0.0 RCVD_IN_MSPIKE_H2 RBL: Average reputation (+2)
[209.85.128.49 listed in wl.mailspike.net]
X-Headers-End: 1vmtu0-0002Ab-6a
Subject: [Openvpn-devel] [PATCH ovpn net] ovpn: detach TCP socket before
invoking close
X-BeenThere: openvpn-devel@lists.sourceforge.net
X-Mailman-Version: 2.1.21
Precedence: list
List-Id: <openvpn-devel.lists.sourceforge.net>
List-Unsubscribe: <https://lists.sourceforge.net/lists/options/openvpn-devel>,
<mailto:openvpn-devel-request@lists.sourceforge.net?subject=unsubscribe>
List-Archive:
<http://sourceforge.net/mailarchive/forum.php?forum_name=openvpn-devel>
List-Post: <mailto:openvpn-devel@lists.sourceforge.net>
List-Help: <mailto:openvpn-devel-request@lists.sourceforge.net?subject=help>
List-Subscribe: <https://lists.sourceforge.net/lists/listinfo/openvpn-devel>,
<mailto:openvpn-devel-request@lists.sourceforge.net?subject=subscribe>
Cc: Sabrina Dubroca <sd@queasysnail.net>
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Errors-To: openvpn-devel-bounces@lists.sourceforge.net
X-getmail-retrieved-from-mailbox: Inbox
X-GMAIL-THRID: =?utf-8?q?1856020028422940463?=
X-GMAIL-MSGID: =?utf-8?q?1856020028422940463?=
|
| Series |
[Openvpn-devel,ovpn,net] ovpn: detach TCP socket before invoking close
|
|
Commit Message
Ralf Lici
Feb. 2, 2026, 1:23 p.m. UTC
Currently, when a connected peer expires (no packets received within the
keepalive interval), we remove the peer and notify userspace of the
deletion. We then insert the peer in the release list and proceed to
detach and release the socket and the peer. This can be problematic with
TCP because, as soon as we send the notification, openvpn will close the
peer's socket and if ovpn_tcp_close is invoked before
ovpn_tcp_socket_detach we incurr in a NULL pointer dereference when
trying to access sk->sk_socket.
Enforce correct ordering by calling ovpn_sock_release before invoking
the original socket close callback. This avoids potential race
conditions and guarantees that we completely detach from the socket once
userspace issues the close command.
Signed-off-by: Ralf Lici <ralf@mandelbit.com>
---
drivers/net/ovpn/tcp.c | 1 +
1 file changed, 1 insertion(+)
Comments
This patch is actually fixing the same issue reported here: https://lore.kernel.org/netdev/176996279620.3109699.15382994681575380467@eldamar.lan/ Cheers, On 02/02/2026 14:23, Ralf Lici wrote: > Currently, when a connected peer expires (no packets received within the > keepalive interval), we remove the peer and notify userspace of the > deletion. We then insert the peer in the release list and proceed to > detach and release the socket and the peer. This can be problematic with > TCP because, as soon as we send the notification, openvpn will close the > peer's socket and if ovpn_tcp_close is invoked before > ovpn_tcp_socket_detach we incurr in a NULL pointer dereference when > trying to access sk->sk_socket. > > Enforce correct ordering by calling ovpn_sock_release before invoking > the original socket close callback. This avoids potential race > conditions and guarantees that we completely detach from the socket once > userspace issues the close command. > > Signed-off-by: Ralf Lici <ralf@mandelbit.com> > --- > drivers/net/ovpn/tcp.c | 1 + > 1 file changed, 1 insertion(+) > > diff --git a/drivers/net/ovpn/tcp.c b/drivers/net/ovpn/tcp.c > index 0d7f30360d87..13d2a8069695 100644 > --- a/drivers/net/ovpn/tcp.c > +++ b/drivers/net/ovpn/tcp.c > @@ -553,6 +553,7 @@ static void ovpn_tcp_close(struct sock *sk, long timeout) > rcu_read_unlock(); > > ovpn_peer_del(sock->peer, OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); > + ovpn_socket_release(peer); > peer->tcp.sk_cb.prot->close(sk, timeout); > ovpn_peer_put(peer); > }
On Mon, 2026-02-02 at 14:23 +0100, Ralf Lici wrote: > Currently, when a connected peer expires (no packets received within > the > keepalive interval), we remove the peer and notify userspace of the > deletion. We then insert the peer in the release list and proceed to > detach and release the socket and the peer. This can be problematic > with > TCP because, as soon as we send the notification, openvpn will close > the > peer's socket and if ovpn_tcp_close is invoked before > ovpn_tcp_socket_detach we incurr in a NULL pointer dereference when > trying to access sk->sk_socket. > > Enforce correct ordering by calling ovpn_sock_release before invoking > the original socket close callback. This avoids potential race > conditions and guarantees that we completely detach from the socket > once > userspace issues the close command. > Fixes: 11851cbd60ea ("ovpn: implement TCP transport") > Signed-off-by: Ralf Lici <ralf@mandelbit.com> > --- > drivers/net/ovpn/tcp.c | 1 + > 1 file changed, 1 insertion(+) > > diff --git a/drivers/net/ovpn/tcp.c b/drivers/net/ovpn/tcp.c > index 0d7f30360d87..13d2a8069695 100644 > --- a/drivers/net/ovpn/tcp.c > +++ b/drivers/net/ovpn/tcp.c > @@ -553,6 +553,7 @@ static void ovpn_tcp_close(struct sock *sk, long > timeout) > rcu_read_unlock(); > > ovpn_peer_del(sock->peer, > OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); > + ovpn_socket_release(peer); > peer->tcp.sk_cb.prot->close(sk, timeout); > ovpn_peer_put(peer); > }
2026-02-02, 14:33:47 +0100, Antonio Quartulli wrote: > This patch is actually fixing the same issue reported here: > > https://lore.kernel.org/netdev/176996279620.3109699.15382994681575380467@eldamar.lan/ So it would be nice to add a Link: <url> tag just above the sign-off/fixes tag (to the netdev post or to the debian bugtracker), and there could also be a Reported-by tag (you may want to check with the reporter if they agree, but the debian bugtracker publicly lists the name and email so I don't see why not). > Cheers, > > On 02/02/2026 14:23, Ralf Lici wrote: > > Currently, when a connected peer expires (no packets received within the > > keepalive interval), we remove the peer and notify userspace of the > > deletion. We then insert the peer in the release list and proceed to > > detach and release the socket and the peer. This can be problematic with > > TCP because, as soon as we send the notification, openvpn will close the > > peer's socket and if ovpn_tcp_close is invoked before > > ovpn_tcp_socket_detach we incurr in a NULL pointer dereference when > > trying to access sk->sk_socket. > > > > Enforce correct ordering by calling ovpn_sock_release before invoking > > the original socket close callback. This avoids potential race > > conditions and guarantees that we completely detach from the socket once > > userspace issues the close command. The comment for ovpn_sock_release says: * This function is expected to be invoked exactly once per peer But it seems in this case we won't be obeying this assumption? Once from ovpn_tcp_close, and again from unlock_ovpn in the timeout handler? IIRC this "invoked exactly once" was the condition that made rcu_replace_pointer(peer->sock, NULL, true); valid, but what will prevent both sides from seeing peer->sock != NULL now? (not really related, but maybe unlock_ovpn should be using llist_for_each_entry_safe since the current peer might go away before we grab the next list item?)
On 06/02/2026 19:43, Sabrina Dubroca wrote: > 2026-02-02, 14:33:47 +0100, Antonio Quartulli wrote: >> This patch is actually fixing the same issue reported here: >> >> https://lore.kernel.org/netdev/176996279620.3109699.15382994681575380467@eldamar.lan/ > > So it would be nice to add a > > Link: <url> > > tag just above the sign-off/fixes tag (to the netdev post or to the > debian bugtracker), and there could also be a Reported-by tag (you may > want to check with the reporter if they agree, but the debian > bugtracker publicly lists the name and email so I don't see why not). > The bug was discovered[1] and the patch was created before we got the Debian report. Hence there is no mention of them. [1]https://github.com/OpenVPN/ovpn-net-next/issues/29 But we can still add the Link tag at least, to make it clear that the ticket is related. >> On 02/02/2026 14:23, Ralf Lici wrote: >>> Currently, when a connected peer expires (no packets received within the >>> keepalive interval), we remove the peer and notify userspace of the >>> deletion. We then insert the peer in the release list and proceed to >>> detach and release the socket and the peer. This can be problematic with >>> TCP because, as soon as we send the notification, openvpn will close the >>> peer's socket and if ovpn_tcp_close is invoked before >>> ovpn_tcp_socket_detach we incurr in a NULL pointer dereference when >>> trying to access sk->sk_socket. >>> >>> Enforce correct ordering by calling ovpn_sock_release before invoking >>> the original socket close callback. This avoids potential race >>> conditions and guarantees that we completely detach from the socket once >>> userspace issues the close command. > > The comment for ovpn_sock_release says: > > * This function is expected to be invoked exactly once per peer > > But it seems in this case we won't be obeying this assumption? Once > from ovpn_tcp_close, and again from unlock_ovpn in the timeout > handler? > > IIRC this "invoked exactly once" was the condition that made > > rcu_replace_pointer(peer->sock, NULL, true); > > valid, but what will prevent both sides from seeing peer->sock != NULL > now? Mh indeed you are right. The rcu_replace_pointer call must be protected now. We have to double check which lock we can safely hold without incurring in any conflict with the rest of the socket lifecycle. Thanks for pointing this out. > > > (not really related, but maybe unlock_ovpn should be using > llist_for_each_entry_safe since the current peer might go away before > we grab the next list item?) hm it will go away at the next RCU cycle, when the RCU-scheduled call to ovpn_peer_release_rcu() will kick in. And indeed we are not protected against that, so yeah, it can theoretically go away. Thanks for spotting this! Wanna send a patch? Otherwise we'll take care of it :) Regards,
On 09/02/2026 11:21, Antonio Quartulli wrote: >>> On 02/02/2026 14:23, Ralf Lici wrote: >>>> Currently, when a connected peer expires (no packets received within >>>> the >>>> keepalive interval), we remove the peer and notify userspace of the >>>> deletion. We then insert the peer in the release list and proceed to >>>> detach and release the socket and the peer. This can be problematic >>>> with >>>> TCP because, as soon as we send the notification, openvpn will close >>>> the >>>> peer's socket and if ovpn_tcp_close is invoked before >>>> ovpn_tcp_socket_detach we incurr in a NULL pointer dereference when >>>> trying to access sk->sk_socket. >>>> >>>> Enforce correct ordering by calling ovpn_sock_release before invoking >>>> the original socket close callback. This avoids potential race >>>> conditions and guarantees that we completely detach from the socket >>>> once >>>> userspace issues the close command. >> >> The comment for ovpn_sock_release says: >> >> * This function is expected to be invoked exactly once per peer >> >> But it seems in this case we won't be obeying this assumption? Once >> from ovpn_tcp_close, and again from unlock_ovpn in the timeout >> handler? >> >> IIRC this "invoked exactly once" was the condition that made >> >> rcu_replace_pointer(peer->sock, NULL, true); >> >> valid, but what will prevent both sides from seeing peer->sock != NULL >> now? > > Mh indeed you are right. > > The rcu_replace_pointer call must be protected now. > We have to double check which lock we can safely hold without incurring > in any conflict with the rest of the socket lifecycle. > > Thanks for pointing this out. What if we go the other way around: we have ovpn_tcp_close() check if the peer was already gone from the ovpn hash or not. If it did, we know the peer is going through its own lifecycle and we skip the ovpn_peer_del() call. This is what I have mind (not tested yet! only compiled): diff --git a/drivers/net/ovpn/tcp.c b/drivers/net/ovpn/tcp.c index 1c02cc06f21b..83e7a5354266 100644 --- a/drivers/net/ovpn/tcp.c +++ b/drivers/net/ovpn/tcp.c @@ -543,8 +543,8 @@ int ovpn_tcp_socket_attach(struct ovpn_socket *ovpn_sock, static void ovpn_tcp_close(struct sock *sk, long timeout) { + struct ovpn_peer *peer, *tmp; struct ovpn_socket *sock; - struct ovpn_peer *peer; rcu_read_lock(); sock = rcu_dereference_sk_user_data(sk); @@ -555,7 +555,26 @@ static void ovpn_tcp_close(struct sock *sk, long timeout) peer = sock->peer; rcu_read_unlock(); - ovpn_peer_del(sock->peer, OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); + spin_lock_bh(&peer->ovpn->lock); + /* if the peer is already unhashed, it means it already + * went through ovpn_peer_remove() due to the actual removal event. + * This tcp_close call is just the socket being cleaned up by + * userspace and we don't need to double process the deletion + * + * NOTE: we need to lock (instead of just rcu-read-lock) because + * we have to synchronize with other calls to ovpn_peer_remove() + * happening with ovpn->lock held. + * This way we are guaranteed to call ovpn_peer_del() without + * racing with another ovpn_peer_remove(). + */ + tmp = ovpn_peer_get_by_id(peer->ovpn, peer->id); + if (tmp) { + ovpn_peer_del(sock->peer, + OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); + ovpn_peer_put(tmp); + } + spin_unlock_bh(&peer->ovpn->lock); + peer->tcp.sk_cb.prot->close(sk, timeout); ovpn_peer_put(peer); } What do you think?
On 09/02/2026 16:32, Antonio Quartulli wrote: > What if we go the other way around: we have ovpn_tcp_close() check if > the peer was already gone from the ovpn hash or not. > > If it did, we know the peer is going through its own lifecycle and we > skip the ovpn_peer_del() call. > > This is what I have mind (not tested yet! only compiled): > > diff --git a/drivers/net/ovpn/tcp.c b/drivers/net/ovpn/tcp.c > index 1c02cc06f21b..83e7a5354266 100644 > --- a/drivers/net/ovpn/tcp.c > +++ b/drivers/net/ovpn/tcp.c > @@ -543,8 +543,8 @@ int ovpn_tcp_socket_attach(struct ovpn_socket > *ovpn_sock, > > static void ovpn_tcp_close(struct sock *sk, long timeout) > { > + struct ovpn_peer *peer, *tmp; > struct ovpn_socket *sock; > - struct ovpn_peer *peer; > > rcu_read_lock(); > sock = rcu_dereference_sk_user_data(sk); > @@ -555,7 +555,26 @@ static void ovpn_tcp_close(struct sock *sk, long > timeout) > peer = sock->peer; > rcu_read_unlock(); > > - ovpn_peer_del(sock->peer, > OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); > + spin_lock_bh(&peer->ovpn->lock); > + /* if the peer is already unhashed, it means it already > + * went through ovpn_peer_remove() due to the actual removal event. > + * This tcp_close call is just the socket being cleaned up by > + * userspace and we don't need to double process the deletion > + * > + * NOTE: we need to lock (instead of just rcu-read-lock) because > + * we have to synchronize with other calls to ovpn_peer_remove() > + * happening with ovpn->lock held. > + * This way we are guaranteed to call ovpn_peer_del() without > + * racing with another ovpn_peer_remove(). > + */ > + tmp = ovpn_peer_get_by_id(peer->ovpn, peer->id); > + if (tmp) { or even better, we could check for !hlist_unhashed(&peer->hash_entry_id), to avoid hitting another peer that just got the same ID re-assigned (nearly impossible, but you know ..) Cheers, > + ovpn_peer_del(sock->peer, > + OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); > + ovpn_peer_put(tmp); > + } > + spin_unlock_bh(&peer->ovpn->lock); > + > peer->tcp.sk_cb.prot->close(sk, timeout); > ovpn_peer_put(peer); > } > > > What do you think? > >
2026-02-09, 16:36:55 +0100, Antonio Quartulli wrote: > On 09/02/2026 16:32, Antonio Quartulli wrote: > > What if we go the other way around: we have ovpn_tcp_close() check if > > the peer was already gone from the ovpn hash or not. > > > > If it did, we know the peer is going through its own lifecycle and we > > skip the ovpn_peer_del() call. To make sure I understand: the bug we're trying to fix here is that we end up calling ovpn_tcp_close() too early, so by the time ovpn_tcp_socket_detach() is called, sk->sk_socket is already NULL so "sk->sk_socket->ops = ..." crashes the kernel. Ralf's patch makes sure ovpn_socket_release() -> ... -> ovpn_tcp_socket_detach() has been called by the time we're done with close(). Is that correct? Then I'm confused by how this would solve the issue. We could still be done with ovpn_tcp_close() before ovpn_peer_keepalive_work() -> unlock_ovpn -> ovpn_socket_release has completed? > > This is what I have mind (not tested yet! only compiled): > > > > diff --git a/drivers/net/ovpn/tcp.c b/drivers/net/ovpn/tcp.c > > index 1c02cc06f21b..83e7a5354266 100644 > > --- a/drivers/net/ovpn/tcp.c > > +++ b/drivers/net/ovpn/tcp.c > > @@ -543,8 +543,8 @@ int ovpn_tcp_socket_attach(struct ovpn_socket > > *ovpn_sock, > > > > static void ovpn_tcp_close(struct sock *sk, long timeout) > > { > > + struct ovpn_peer *peer, *tmp; > > struct ovpn_socket *sock; > > - struct ovpn_peer *peer; > > > > rcu_read_lock(); > > sock = rcu_dereference_sk_user_data(sk); > > @@ -555,7 +555,26 @@ static void ovpn_tcp_close(struct sock *sk, long > > timeout) > > peer = sock->peer; > > rcu_read_unlock(); > > > > - ovpn_peer_del(sock->peer, > > OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); > > + spin_lock_bh(&peer->ovpn->lock); > > + /* if the peer is already unhashed, it means it already > > + * went through ovpn_peer_remove() due to the actual removal event. > > + * This tcp_close call is just the socket being cleaned up by > > + * userspace and we don't need to double process the deletion > > + * > > + * NOTE: we need to lock (instead of just rcu-read-lock) because > > + * we have to synchronize with other calls to ovpn_peer_remove() > > + * happening with ovpn->lock held. > > + * This way we are guaranteed to call ovpn_peer_del() without > > + * racing with another ovpn_peer_remove(). > > + */ > > + tmp = ovpn_peer_get_by_id(peer->ovpn, peer->id); > > + if (tmp) { > > or even better, we could check for !hlist_unhashed(&peer->hash_entry_id), to But hash_entry_id isn't used in P2P mode? > avoid hitting another peer that just got the same ID re-assigned (nearly > impossible, but you know ..) > > > Cheers, > > > + ovpn_peer_del(sock->peer, ovpn_peer_del wants to acquire peer->ovpn->lock which we just took, you'd have to rework that function. > > + OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); > > + ovpn_peer_put(tmp); > > + } > > + spin_unlock_bh(&peer->ovpn->lock); > > + > > peer->tcp.sk_cb.prot->close(sk, timeout); > > ovpn_peer_put(peer); > > }
On 09/02/2026 17:34, Sabrina Dubroca wrote: > 2026-02-09, 16:36:55 +0100, Antonio Quartulli wrote: >> On 09/02/2026 16:32, Antonio Quartulli wrote: >>> What if we go the other way around: we have ovpn_tcp_close() check if >>> the peer was already gone from the ovpn hash or not. >>> >>> If it did, we know the peer is going through its own lifecycle and we >>> skip the ovpn_peer_del() call. > > To make sure I understand: the bug we're trying to fix here is that we > end up calling ovpn_tcp_close() too early, so by the time > ovpn_tcp_socket_detach() is called, sk->sk_socket is already NULL so > "sk->sk_socket->ops = ..." crashes the kernel. Ralf's patch makes sure > ovpn_socket_release() -> ... -> ovpn_tcp_socket_detach() has been > called by the time we're done with close(). Is that correct? > > Then I'm confused by how this would solve the issue. We could still be > done with ovpn_tcp_close() before ovpn_peer_keepalive_work() -> > unlock_ovpn -> ovpn_socket_release has completed? Mh ok. At first I wrote down my entire theory, then I realized it was flawed and erased it again. You're right, we still have tcp_close() being invoked before ovpn_socket_release(), which is what is causing the issue due to sk->sk_socket == NULL. Do you think we could have a reliable way to check in ovpn_socket_release() is sk->sk_socket is NULL and, if so, just bail out? That'd be telling us: "the socket is already gone - no need to continue and crash". Cheers, > >>> This is what I have mind (not tested yet! only compiled): >>> >>> diff --git a/drivers/net/ovpn/tcp.c b/drivers/net/ovpn/tcp.c >>> index 1c02cc06f21b..83e7a5354266 100644 >>> --- a/drivers/net/ovpn/tcp.c >>> +++ b/drivers/net/ovpn/tcp.c >>> @@ -543,8 +543,8 @@ int ovpn_tcp_socket_attach(struct ovpn_socket >>> *ovpn_sock, >>> >>> static void ovpn_tcp_close(struct sock *sk, long timeout) >>> { >>> + struct ovpn_peer *peer, *tmp; >>> struct ovpn_socket *sock; >>> - struct ovpn_peer *peer; >>> >>> rcu_read_lock(); >>> sock = rcu_dereference_sk_user_data(sk); >>> @@ -555,7 +555,26 @@ static void ovpn_tcp_close(struct sock *sk, long >>> timeout) >>> peer = sock->peer; >>> rcu_read_unlock(); >>> >>> - ovpn_peer_del(sock->peer, >>> OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); >>> + spin_lock_bh(&peer->ovpn->lock); >>> + /* if the peer is already unhashed, it means it already >>> + * went through ovpn_peer_remove() due to the actual removal event. >>> + * This tcp_close call is just the socket being cleaned up by >>> + * userspace and we don't need to double process the deletion >>> + * >>> + * NOTE: we need to lock (instead of just rcu-read-lock) because >>> + * we have to synchronize with other calls to ovpn_peer_remove() >>> + * happening with ovpn->lock held. >>> + * This way we are guaranteed to call ovpn_peer_del() without >>> + * racing with another ovpn_peer_remove(). >>> + */ >>> + tmp = ovpn_peer_get_by_id(peer->ovpn, peer->id); >>> + if (tmp) { >> >> or even better, we could check for !hlist_unhashed(&peer->hash_entry_id), to > > But hash_entry_id isn't used in P2P mode? > >> avoid hitting another peer that just got the same ID re-assigned (nearly >> impossible, but you know ..) >> >> >> Cheers, >> >>> + ovpn_peer_del(sock->peer, > > ovpn_peer_del wants to acquire peer->ovpn->lock which we just took, > you'd have to rework that function. > >>> + OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); >>> + ovpn_peer_put(tmp); >>> + } >>> + spin_unlock_bh(&peer->ovpn->lock); >>> + >>> peer->tcp.sk_cb.prot->close(sk, timeout); >>> ovpn_peer_put(peer); >>> } >
On 10/02/2026 01:07, Antonio Quartulli wrote: > On 09/02/2026 17:34, Sabrina Dubroca wrote: >> 2026-02-09, 16:36:55 +0100, Antonio Quartulli wrote: >>> On 09/02/2026 16:32, Antonio Quartulli wrote: >>>> What if we go the other way around: we have ovpn_tcp_close() check if >>>> the peer was already gone from the ovpn hash or not. >>>> >>>> If it did, we know the peer is going through its own lifecycle and we >>>> skip the ovpn_peer_del() call. >> >> To make sure I understand: the bug we're trying to fix here is that we >> end up calling ovpn_tcp_close() too early, so by the time >> ovpn_tcp_socket_detach() is called, sk->sk_socket is already NULL so >> "sk->sk_socket->ops = ..." crashes the kernel. Ralf's patch makes sure >> ovpn_socket_release() -> ... -> ovpn_tcp_socket_detach() has been >> called by the time we're done with close(). Is that correct? >> >> Then I'm confused by how this would solve the issue. We could still be >> done with ovpn_tcp_close() before ovpn_peer_keepalive_work() -> >> unlock_ovpn -> ovpn_socket_release has completed? > As we have already stated, the crux of the issue is the moment tcp_close() sets sk->sk_socket to NULL. This happens in: 1. tcp_close(sk, timeout); 2. __tcp_close(sk, timeout); 3. sock_orphan(sk); 4. sk_set_socket(sk, NULL); When we later/concurrently invoke ovpn_tcp_socket_release() we dereference sk->sk_socket, which is now NULL. sock_orphan()is actually pretty small: 2123 static inline void sock_orphan(struct sock *sk) 2124 { 2125 write_lock_bh(&sk->sk_callback_lock); 2126 sock_set_flag(sk, SOCK_DEAD); 2127 sk_set_socket(sk, NULL); 2128 sk->sk_wq = NULL; 2129 write_unlock_bh(&sk->sk_callback_lock); 2130 } and as we can see it holds and then releases sk_callback_lock. I remember that we were originally using this lock when setting the callbacks, but we eventually dropped it (around v3 or v4 of the original patchset, but I couldn't find the reason). Maybe we thought there was no need to hold that lock? Now, if we actually acquire this lock in ovpn_tcp_socket_detach(), check that sk_socket is non-NULL, restores the CB and release it, wouldn't we solve our problem? Cheers,
2026-02-10, 14:50:47 +0100, Antonio Quartulli wrote: > On 10/02/2026 01:07, Antonio Quartulli wrote: > > On 09/02/2026 17:34, Sabrina Dubroca wrote: > > > 2026-02-09, 16:36:55 +0100, Antonio Quartulli wrote: > > > > On 09/02/2026 16:32, Antonio Quartulli wrote: > > > > > What if we go the other way around: we have ovpn_tcp_close() check if > > > > > the peer was already gone from the ovpn hash or not. > > > > > > > > > > If it did, we know the peer is going through its own lifecycle and we > > > > > skip the ovpn_peer_del() call. > > > > > > To make sure I understand: the bug we're trying to fix here is that we > > > end up calling ovpn_tcp_close() too early, so by the time > > > ovpn_tcp_socket_detach() is called, sk->sk_socket is already NULL so > > > "sk->sk_socket->ops = ..." crashes the kernel. Ralf's patch makes sure > > > ovpn_socket_release() -> ... -> ovpn_tcp_socket_detach() has been > > > called by the time we're done with close(). Is that correct? > > > > > > Then I'm confused by how this would solve the issue. We could still be > > > done with ovpn_tcp_close() before ovpn_peer_keepalive_work() -> > > > unlock_ovpn -> ovpn_socket_release has completed? > > > Mh ok. At first I wrote down my entire theory, then I realized it > was flawed and erased it again. I've done that as well at least once in this thread :) > As we have already stated, the crux of the issue is the moment tcp_close() > sets sk->sk_socket to NULL. > > This happens in: > 1. tcp_close(sk, timeout); > 2. __tcp_close(sk, timeout); > 3. sock_orphan(sk); > 4. sk_set_socket(sk, NULL); > > When we later/concurrently invoke ovpn_tcp_socket_release() we dereference > sk->sk_socket, which is now NULL. > > sock_orphan()is actually pretty small: > > 2123 static inline void sock_orphan(struct sock *sk) > 2124 { > 2125 write_lock_bh(&sk->sk_callback_lock); > 2126 sock_set_flag(sk, SOCK_DEAD); > 2127 sk_set_socket(sk, NULL); > 2128 sk->sk_wq = NULL; > 2129 write_unlock_bh(&sk->sk_callback_lock); > 2130 } > > and as we can see it holds and then releases sk_callback_lock. > > I remember that we were originally using this lock when setting the > callbacks, but we eventually dropped it (around v3 or v4 of the original > patchset, but I couldn't find the reason). > Maybe we thought there was no need to hold that lock? Possibly :) Or holding a spinlock at that stage was no longer possible due to some changes (maybe we'd have done sleeping ops under it), but either way, the locking probably changed another 5 times afterwards :) > Now, if we actually acquire this lock in ovpn_tcp_socket_detach(), check > that sk_socket is non-NULL, restores the CB and release it, wouldn't we > solve our problem? I think it would, but maybe post this to netdev with Jakub/Paolo/etc in CC for confirmation (answers may be slow due to the merge window). Expanding the commit message to describe a bit better the order of events leading to the crash would be good too (maybe also summarize this thread, ie explain other idea(s) and why they don't work). Otherwise we might have to come up with a scheme that delays ovpn_tcp_close() until the pending/in progress ovpn_socket_release() has completed. Let's hope not... (I was looking into another idea, making ovpn_socket_release's rcu_replace_pointer atomic by using xchg instead, but this could still end up calling tcp_close before detach() is done if ovpn_tcp_close starts immediately after we've swapped the pointer)
diff --git a/drivers/net/ovpn/tcp.c b/drivers/net/ovpn/tcp.c index 0d7f30360d87..13d2a8069695 100644 --- a/drivers/net/ovpn/tcp.c +++ b/drivers/net/ovpn/tcp.c @@ -553,6 +553,7 @@ static void ovpn_tcp_close(struct sock *sk, long timeout) rcu_read_unlock(); ovpn_peer_del(sock->peer, OVPN_DEL_PEER_REASON_TRANSPORT_DISCONNECT); + ovpn_socket_release(peer); peer->tcp.sk_cb.prot->close(sk, timeout); ovpn_peer_put(peer); }