mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-05 14:02:00 +02:00
* feat(nfs): add UDP MOUNT v3 responder
The upstream willscott/go-nfs library only serves the MOUNT protocol
over TCP. Linux's mount.nfs and the in-kernel NFS client default
mountproto to UDP in many configurations, so against a stock weed nfs
deployment the kernel queries portmap for "MOUNT v3 UDP", gets port=0
("not registered"), and either falls back inconsistently or surfaces
EPROTONOSUPPORT — surfacing as the user-visible "requested NFS version
or transport protocol is not supported" reported in #9263. The user has
to add `mountproto=tcp` or `mountport=2049` to mount options to coerce
TCP just for the MOUNT phase.
Add a small UDP responder that speaks just enough of MOUNT v3 to handle
the procedures the kernel actually invokes during mount setup and
teardown: NULL, MNT, and UMNT. The wire layout for MNT mirrors
handler.go's TCP path so both transports produce the same root
filehandle and the same auth flavor list for the same export. Other
v3 procedures (DUMP, EXPORT, UMNTALL) cleanly return PROC_UNAVAIL.
This commit only adds the responder; portmap-advertise and Server.Start
wire-up follow in subsequent commits so each step stays independently
reviewable.
References: RFC 1813 §5 (NFSv3/MOUNTv3), RFC 5531 (RPC). Existing
constants and parseRPCCall / encodeAcceptedReply helpers from
portmap.go are reused so behaviour stays consistent across both UDP
listening goroutines.
* feat(nfs): advertise UDP MOUNT v3 in the portmap responder
The portmap responder advertised TCP-only entries because go-nfs only
serves TCP, but with the new UDP MOUNT responder in place we can now
honestly advertise MOUNT v3 over UDP as well. Linux clients whose
default mountproto is UDP query portmap during mount setup; if the
answer is "not registered" some kernels translate the result to
EPROTONOSUPPORT instead of falling back to TCP, which is exactly the
failure pattern reported in #9263.
Add the entry, refresh the doc comment, and extend the existing
GETPORT and DUMP unit tests so a regression that drops the entry shows
up at unit-test granularity rather than only in an end-to-end mount.
* feat(nfs): start UDP MOUNT v3 responder alongside the TCP NFS listener
Plug the new mountUDPServer into Server.Start so it comes up on the
same bind/port as the TCP NFS listener. Started before portmap so a
portmap query that races a fast client never returns a UDP MOUNT entry
the responder isn't actually answering, and shut down via the same
defer chain so a portmap-or-listener startup failure doesn't leave the
UDP responder dangling.
The portmap startup log now reflects all three advertised entries
(NFS v3 tcp, MOUNT v3 tcp, MOUNT v3 udp) so operators can confirm at a
glance that the UDP MOUNT path is up.
Verified end-to-end: built a Linux/arm64 binary, ran weed nfs in a
container with -portmap.bind, and mounted from another container using
both the user-reported failing setup from #9263 (vers=3 + tcp without
mountport) and an explicit mountproto=udp to force the new code path.
The trace `mount.nfs: trying ... prog 100005 vers 3 prot UDP port 2049`
now leads to a successful mount instead of EPROTONOSUPPORT.
* docs(nfs): note that the plain mount form works on UDP-default clients
With UDP MOUNT v3 now served alongside TCP, the only path that ever
required mountproto=tcp / mountport=2049 — clients whose default
mountproto is UDP — works against the plain mount example. Update the
startup mount hint and the `weed nfs` long help so users don't go
hunting for a mount-option workaround that no longer applies.
The "without -portmap.bind" branch is unchanged: that path still has
to bypass portmap entirely because there is no portmap responder for
the kernel to query.
* test(nfs): add kernel-mount e2e tests under test/nfs
The existing test/nfs/ harness boots a real master + volume + filer +
weed nfs subprocess stack and drives it via go-nfs-client. That covers
protocol behaviour from a Go client's perspective, but anything
mis-coded once a real Linux kernel parses the wire bytes is invisible:
both ends of the test use the same RPC library, so identical bugs
round-trip cleanly. The two NFS issues hit recently were exactly that
shape — NFSv4 mis-routed to v3 SETATTR (#9262) and missing UDP MOUNT v3
— and only surfaced in a real client.
Add three end-to-end tests that mount the harness's running NFS server
through the in-tree Linux client:
- TestKernelMountV3TCP: NFSv3 + MOUNT v3 over TCP (baseline).
- TestKernelMountV3MountProtoUDP: NFSv3 over TCP, MOUNT v3 over UDP
only — regression test for the new UDP MOUNT v3 responder.
- TestKernelMountV4RejectsCleanly: vers=4 against the v3-only server,
asserting the kernel surfaces a protocol/version-level error rather
than a generic "mount system call failed" — regression test for the
PROG_MISMATCH path from #9262.
The tests pass explicit port=/mountport= mount options so the kernel
never queries portmap, which means the harness doesn't need to bind
the privileged port 111 and won't collide with a system rpcbind on a
shared CI runner. They t.Skip cleanly when the host isn't Linux, when
mount.nfs isn't installed, or when the test process isn't running as
root.
Run locally with:
cd test/nfs
sudo go test -v -run TestKernelMount ./...
CI wiring follows in the next commit.
* ci(nfs): run kernel-mount e2e tests in nfs-tests workflow
Wire the new TestKernelMount* tests from test/nfs into the existing
NFS workflow:
- Existing protocol-layer step now skips '^TestKernelMount' so a
"skipped because not root" line doesn't appear on every run.
- New "Install kernel NFS client" step pulls nfs-common (mount.nfs +
helpers) and netbase (/etc/protocols, which mount.nfs's protocol-
name lookups need to resolve `tcp`/`udp`).
- New privileged step runs only the kernel-mount tests under sudo,
preserving PATH and pointing GOMODCACHE/GOCACHE at the user's
caches so the second `go test` invocation reuses already-built
test binaries instead of redownloading modules under root.
The summary block now lists the three kernel-mount cases explicitly
so a regression on either of #9262 or this PR's UDP MOUNT change is
traceable from the workflow run page.
266 lines
8.4 KiB
Go
266 lines
8.4 KiB
Go
package nfs
|
|
|
|
import (
|
|
"encoding/binary"
|
|
"fmt"
|
|
"net"
|
|
"sync"
|
|
"time"
|
|
|
|
"github.com/seaweedfs/seaweedfs/weed/filer"
|
|
"github.com/seaweedfs/seaweedfs/weed/glog"
|
|
"github.com/seaweedfs/seaweedfs/weed/util"
|
|
)
|
|
|
|
// The upstream willscott/go-nfs library only serves the MOUNT protocol over
|
|
// TCP. Linux's mount.nfs and the in-kernel NFS client default `mountproto` to
|
|
// UDP in many configurations, so against a stock `weed nfs` deployment the
|
|
// kernel queries portmap for "MOUNT v3 UDP", gets port=0 ("not registered"),
|
|
// and either falls back inconsistently or surfaces EPROTONOSUPPORT
|
|
// ("requested NFS version or transport protocol is not supported"). The user
|
|
// either has to add `mountproto=tcp` / `mountport=2049` to their mount
|
|
// options or guess that their distro happens to fall back to TCP on its own.
|
|
//
|
|
// This responder closes that gap. It speaks just enough of MOUNT v3 to handle
|
|
// MOUNT_NULL / MOUNT_MNT / MOUNT_UMNT over UDP — the only procedures the
|
|
// kernel actually invokes during mount setup and teardown — so plain
|
|
// `mount -t nfs <host>:<export> /mnt` works without any client-side protocol
|
|
// hints. The protocol layout is intentionally identical to the TCP MOUNT
|
|
// handler in handler.go's Mount() so the two paths return the same
|
|
// filehandle and the same set of auth flavors for the same export.
|
|
//
|
|
// References: RFC 1813 §5 (NFSv3/MOUNTv3), RFC 5531 (RPC).
|
|
|
|
const (
|
|
mountUDPMaxRecord = 32 * 1024
|
|
|
|
// mountUDPRetryBackoff mirrors portmapRetryBackoff so the two
|
|
// listening goroutines back off identically under host pressure.
|
|
mountUDPRetryBackoff = 50 * time.Millisecond
|
|
|
|
mountVersion = 3
|
|
|
|
mountProcNull = 0
|
|
mountProcMnt = 1
|
|
mountProcUmnt = 3
|
|
|
|
// MOUNT v3 status codes (mountstat3 in RFC 1813 §5.1.1).
|
|
mnt3StatOK uint32 = 0
|
|
mnt3ErrAcces uint32 = 13
|
|
mnt3ErrNoEnt uint32 = 2
|
|
mnt3ErrNotDir uint32 = 20
|
|
|
|
// XDR opaque length cap for dirpath. RFC 1813 §5.1 limits MNTPATHLEN
|
|
// to 1024; cap a bit higher for headroom and reject anything beyond.
|
|
mountUDPMaxPathLen = 4096
|
|
|
|
// AuthFlavor numeric IDs (matches go-nfs and RFC 5531 §8).
|
|
authFlavorNull = 0
|
|
authFlavorUnix = 1
|
|
)
|
|
|
|
// mountUDPServer answers MOUNT v3 RPCs over UDP. It listens on the same port
|
|
// the NFS TCP server uses (2049 by default), since that's what we advertise
|
|
// via portmap, and shares the parent Server's exportRoot, exportID, and
|
|
// client allowlist so the UDP MOUNT path applies the same access policy as
|
|
// the TCP path.
|
|
type mountUDPServer struct {
|
|
bindIP string
|
|
port int
|
|
server *Server
|
|
|
|
udpConn *net.UDPConn
|
|
|
|
mu sync.Mutex
|
|
closed bool
|
|
done chan struct{}
|
|
wg sync.WaitGroup
|
|
}
|
|
|
|
func newMountUDPServer(bindIP string, port int, server *Server) *mountUDPServer {
|
|
return &mountUDPServer{
|
|
bindIP: bindIP,
|
|
port: port,
|
|
server: server,
|
|
done: make(chan struct{}),
|
|
}
|
|
}
|
|
|
|
func (m *mountUDPServer) Start() error {
|
|
addr := net.JoinHostPort(m.bindIP, fmt.Sprintf("%d", m.port))
|
|
udpAddr, err := net.ResolveUDPAddr("udp", addr)
|
|
if err != nil {
|
|
return fmt.Errorf("mount udp resolve %s: %w", addr, err)
|
|
}
|
|
udpConn, err := net.ListenUDP("udp", udpAddr)
|
|
if err != nil {
|
|
return fmt.Errorf("mount udp listen %s: %w", addr, err)
|
|
}
|
|
m.udpConn = udpConn
|
|
m.wg.Add(1)
|
|
go func() {
|
|
defer m.wg.Done()
|
|
m.serve()
|
|
}()
|
|
return nil
|
|
}
|
|
|
|
func (m *mountUDPServer) Close() error {
|
|
m.mu.Lock()
|
|
if m.closed {
|
|
m.mu.Unlock()
|
|
return nil
|
|
}
|
|
m.closed = true
|
|
close(m.done)
|
|
m.mu.Unlock()
|
|
if m.udpConn != nil {
|
|
_ = m.udpConn.Close()
|
|
}
|
|
m.wg.Wait()
|
|
return nil
|
|
}
|
|
|
|
func (m *mountUDPServer) isClosed() bool {
|
|
m.mu.Lock()
|
|
defer m.mu.Unlock()
|
|
return m.closed
|
|
}
|
|
|
|
func (m *mountUDPServer) serve() {
|
|
buf := make([]byte, mountUDPMaxRecord)
|
|
for {
|
|
n, addr, err := m.udpConn.ReadFromUDP(buf)
|
|
if err != nil {
|
|
if m.isClosed() {
|
|
return
|
|
}
|
|
// Transient read failure: log, back off, keep the
|
|
// responder alive — same pattern as portmap UDP.
|
|
glog.V(1).Infof("mount udp read: %v", err)
|
|
select {
|
|
case <-m.done:
|
|
return
|
|
case <-time.After(mountUDPRetryBackoff):
|
|
continue
|
|
}
|
|
}
|
|
// Apply the parent server's client allowlist before we even
|
|
// look at the RPC bytes, mirroring the TCP path's
|
|
// allowlistListener wrapping.
|
|
if m.server != nil && m.server.clientAuthorizer != nil && !m.server.clientAuthorizer.isAllowedAddr(addr) {
|
|
glog.V(1).Infof("mount udp: rejecting unauthorized client %s", addr)
|
|
continue
|
|
}
|
|
reply := m.handleCall(buf[:n], addr)
|
|
if reply == nil {
|
|
continue
|
|
}
|
|
if _, err := m.udpConn.WriteToUDP(reply, addr); err != nil {
|
|
glog.V(1).Infof("mount udp write to %s: %v", addr, err)
|
|
}
|
|
}
|
|
}
|
|
|
|
// handleCall classifies one RPC CALL message and returns the encoded reply,
|
|
// or nil if the call is malformed enough to drop silently.
|
|
func (m *mountUDPServer) handleCall(callBuf []byte, addr *net.UDPAddr) []byte {
|
|
xid, prog, vers, proc, args, err := parseRPCCall(callBuf)
|
|
if err != nil {
|
|
return nil
|
|
}
|
|
if prog != mountProgram {
|
|
return encodeAcceptedReply(xid, rpcAcceptProgUnavail, nil)
|
|
}
|
|
if vers != mountVersion {
|
|
// Mismatch — advertise the v3..v3 we actually support.
|
|
body := make([]byte, 8)
|
|
binary.BigEndian.PutUint32(body[0:4], mountVersion)
|
|
binary.BigEndian.PutUint32(body[4:8], mountVersion)
|
|
return encodeAcceptedReply(xid, rpcAcceptProgMismatch, body)
|
|
}
|
|
|
|
switch proc {
|
|
case mountProcNull:
|
|
return encodeAcceptedReply(xid, rpcAcceptSuccess, nil)
|
|
case mountProcMnt:
|
|
return m.handleMount(xid, args, addr)
|
|
case mountProcUmnt:
|
|
// Stateless server: there's nothing to forget, just acknowledge.
|
|
// The client sends back the dirpath in args; we don't need to
|
|
// validate it here because UMNT has no return data.
|
|
return encodeAcceptedReply(xid, rpcAcceptSuccess, nil)
|
|
default:
|
|
// MOUNT v3 also defines DUMP / EXPORT / UMNTALL but the kernel
|
|
// mount path doesn't invoke them. Returning PROC_UNAVAIL is
|
|
// the protocol-correct response.
|
|
return encodeAcceptedReply(xid, rpcAcceptProcUnavail, nil)
|
|
}
|
|
}
|
|
|
|
// handleMount implements MOUNT v3 MNT. The wire format is RFC 1813 §5.1.4:
|
|
//
|
|
// MOUNT3args { dirpath3 dirpath; } // XDR opaque
|
|
// MOUNT3res { mountstat3 status; if OK { handle, auth_flavors[] } }
|
|
//
|
|
// We mirror handler.go's Mount(): export-path mismatch returns NoEnt; root
|
|
// inode is encoded as a synthetic directory filehandle so it round-trips with
|
|
// the TCP MOUNT path without an extra filer round-trip per UDP MOUNT call.
|
|
func (m *mountUDPServer) handleMount(xid uint32, args []byte, addr *net.UDPAddr) []byte {
|
|
if len(args) < 4 {
|
|
return encodeAcceptedReply(xid, rpcAcceptGarbageArgs, nil)
|
|
}
|
|
pathLen := binary.BigEndian.Uint32(args[0:4])
|
|
if pathLen > mountUDPMaxPathLen {
|
|
return encodeAcceptedReply(xid, rpcAcceptGarbageArgs, nil)
|
|
}
|
|
padded := (pathLen + 3) &^ 3
|
|
if uint32(len(args)) < 4+padded {
|
|
return encodeAcceptedReply(xid, rpcAcceptGarbageArgs, nil)
|
|
}
|
|
dirpath := string(args[4 : 4+pathLen])
|
|
|
|
requestedPath := normalizeExportRoot(util.FullPath(dirpath))
|
|
if requestedPath != m.server.exportRoot {
|
|
glog.V(1).Infof("mount udp: client %s requested %q but export is %q", addr, dirpath, m.server.exportRoot)
|
|
return encodeMountStatus(xid, mnt3ErrNoEnt)
|
|
}
|
|
|
|
rootHandle := NewFileHandle(m.server.exportID, FileHandleKindDirectory, 0, filer.InodeIndexInitialGeneration).Encode()
|
|
flavors := []uint32{authFlavorNull, authFlavorUnix}
|
|
return encodeMountSuccess(xid, rootHandle, flavors)
|
|
}
|
|
|
|
// encodeMountStatus returns a MOUNT MNT reply carrying just an error status.
|
|
// Per RFC 1813 §5.1.4 a non-OK status terminates the response — no handle or
|
|
// flavors follow.
|
|
func encodeMountStatus(xid, status uint32) []byte {
|
|
body := make([]byte, 4)
|
|
binary.BigEndian.PutUint32(body, status)
|
|
return encodeAcceptedReply(xid, rpcAcceptSuccess, body)
|
|
}
|
|
|
|
// encodeMountSuccess builds the OK MOUNT MNT reply: status=OK, file handle
|
|
// (XDR opaque), and the supported auth_flavors list.
|
|
func encodeMountSuccess(xid uint32, handle []byte, flavors []uint32) []byte {
|
|
handleLen := uint32(len(handle))
|
|
handlePadded := (handleLen + 3) &^ 3
|
|
bodyLen := 4 + 4 + handlePadded + 4 + 4*uint32(len(flavors))
|
|
|
|
body := make([]byte, bodyLen)
|
|
binary.BigEndian.PutUint32(body[0:4], mnt3StatOK)
|
|
binary.BigEndian.PutUint32(body[4:8], handleLen)
|
|
copy(body[8:8+handleLen], handle)
|
|
// Trailing pad bytes are already zero from make().
|
|
|
|
pos := 8 + handlePadded
|
|
binary.BigEndian.PutUint32(body[pos:pos+4], uint32(len(flavors)))
|
|
pos += 4
|
|
for _, fl := range flavors {
|
|
binary.BigEndian.PutUint32(body[pos:pos+4], fl)
|
|
pos += 4
|
|
}
|
|
|
|
return encodeAcceptedReply(xid, rpcAcceptSuccess, body)
|
|
}
|