OxideBSD

OxideBSD sockets and local sockets: design specification

Status: accepted design, not yet implemented. Target release: v0.3.0.

The key words MUST, MUST NOT, SHOULD, SHOULD NOT and MAY are to be interpreted as described in RFC 2119. Interfaces are documented in the manual pages socket(2), sendmsg(2), recvmsg(2), getsockopt(2), getpeereid(3) and unix(4); this document records the design. Where FreeBSD, NetBSD and OpenBSD agree it follows them; where they differ, the majority. Linux interfaces are provided alongside where §8 and §9 say so. SYSLOG.md depends on this document (/dev/log).

The source layout is the BSDs’: the socket layer and local sockets are interprocess communication and live in sys/kern (uipc_*), the Internet protocols in sys/netinet, interfaces in sys/net.

1. Scope

1.1. A generic socket layer in the kernel, into which every address family plugs: the existing AF_INET protocols (UDP, TCP, raw ICMP) and the new AF_UNIX.

1.2. AF_UNIX (AF_LOCAL) sockets of types SOCK_STREAM, SOCK_DGRAM and SOCK_SEQPACKET, named in the file system or in the abstract namespace, or unnamed.

1.3. Descriptor passing (SCM_RIGHTS) and credential passing, in the BSD form and the Linux form.

1.4. A socket system-call interface that carries every argument POSIX defines, replacing the reduced forms in use today (§4).

1.5. Out of scope: AF_INET6, a loopback interface, and SO_PASSSEC/SCM_SECURITY.

2. Components

Path Role
sys/kern/uipc_socket.rs The socket object, the protocol switch, and the family-independent system calls
sys/kern/uipc_usrreq.rs AF_UNIX: addresses, connections, buffers, descriptor and credential passing
sys/netinet/ AF_INET protocols (ip, udp, tcp, icmp, arp), as protocol-switch entries
sys/net/ Interfaces and Ethernet
sys/drivers/rtl8139.rs The network interface driver (moved from sys/net/)
sys/modules/socket Registers the socket system calls, for every family (renamed from sys/modules/net)
sys/modules/oxfs Socket inodes (S_IFSOCK)
external/mit/musl src/network/* wrappers, getpeereid(3), credential structures

3. The socket layer

3.1. Every socket is an open file description of kind Socket referring to one kernel socket object. The object records the domain, type and protocol; its state (bound, listening, connecting, connected, read side shut, write side shut); its pending error (SO_ERROR); its options; a receive buffer; and the protocol’s own control block.

3.2. Protocol switch. Each supported (domain, type, protocol) has one entry providing: attach, bind, connect, listen, accept, send, receive, shutdown, readiness, local and peer address, option get and set, and detach. The generic layer MUST NOT contain code specific to a family. socket(2) with an unsupported domain fails with EAFNOSUPPORT, an unsupported type or protocol with EPROTONOSUPPORT (or EPROTOTYPE for a protocol that exists under another type).

3.3. Family-independent behavior implemented once by the generic layer: 1. SOCK_CLOEXEC and SOCK_NONBLOCK in socket(2), socketpair(2) and accept4(2). 2. Blocking, O_NONBLOCK, MSG_DONTWAIT, SO_RCVTIMEO and SO_SNDTIMEO (EAGAIN on expiry). 3. Interruption by a signal: EINTR, or a restart under SA_RESTART when no data has been transferred and no timeout is set, as for any other slow system call. 4. The MSG_* flags: MSG_PEEK, MSG_WAITALL, MSG_DONTWAIT, MSG_TRUNC, MSG_CTRUNC, MSG_EOR, MSG_NOSIGNAL, MSG_CMSG_CLOEXEC. A flag that a protocol does not support (MSG_OOB on AF_UNIX, for example) fails with EOPNOTSUPP; flags are never ignored. 5. The SOL_SOCKET options SO_TYPE, SO_DOMAIN, SO_PROTOCOL, SO_ERROR, SO_ACCEPTCONN, SO_RCVBUF, SO_SNDBUF, SO_RCVLOWAT, SO_RCVTIMEO, SO_SNDTIMEO, SO_REUSEADDR, SO_KEEPALIVE, SO_LINGER, SO_NOSIGPIPE and SO_PEERCRED (§9.2). An unknown option fails with ENOPROTOOPT. 6. SIGPIPE: sending on a stream or sequenced-packet socket whose write side is shut down, or whose peer is gone, fails with EPIPE and sends SIGPIPE to the calling thread, unless MSG_NOSIGNAL or SO_NOSIGPIPE is set. 7. read(2), write(2), readv(2) and writev(2) on a socket are recvmsg/sendmsg with no address, no control data and no flags. 8. fstat(2) reports S_IFSOCK. 9. Readiness for poll(2), ppoll(2) and select(2): readable when data, end-of-file, a pending error or (on a listening socket) a pending connection is available; writable when the send would not block. POLLHUP once both directions are shut down or the peer is gone.

3.4. Waiting. AF_UNIX sockets change state only when another process runs, so a waiter blocks and is woken, as for pipes. AF_INET sockets keep the existing pulled model (the network interface is serviced by the waiter).

Rationale. Today each socket call walks a fixed chain (UDP, then TCP, then ICMP) and the C library drops arguments the system calls cannot carry. A protocol switch is how every BSD kernel structures sockets; it makes a new family a table entry instead of another link in the chain.

4. System-call interface

4.1. The native ABI passes at most four arguments. Calls that need more take a pointer to a structure in the caller’s memory, as the *at() family does.

4.2. The socket system calls are:

Call Number Arguments
socket 140 (domain, type, protocol)
bind 141 (fd, addr, addrlen)
connect 145 (fd, addr, addrlen)
listen 146 (fd, backlog)
accept 147 (fd, addr, addrlen_ptr)
socketpair 149 (domain, type, protocol, sv)
shutdown 152 (fd, how)
getsockname 559 (fd, addr, addrlen_ptr)
sendmsg 577 (fd, msghdr, flags)
recvmsg 578 (fd, msghdr, flags)
getsockopt 579 (fd, sockopt)
setsockopt 580 (fd, sockopt)
getpeername 581 (fd, addr, addrlen_ptr)
accept4 582 (fd, addr, addrlen_ptr, flags)

msghdr is the C library’s struct msghdr. sockopt points to { int64 level; int64 name; uint64 optval; uint64 optlen; }, where optlen is the option’s length for setsockopt and the address of a socklen_t for getsockopt.

4.3. sendto, recvfrom, send and recv MUST be implemented in the C library over sendmsg and recvmsg. Every address length passed in is honored and every address length returned is the actual length of the address (§5.4).

4.4. The former sendto (142), recvfrom (143) and three-argument setsockopt (144) numbers are retired: they MUST fail with ENOSYS and MUST NOT be reassigned.

4.5. Limits: at most IOV_MAX (1024) elements in msg_iov; at most 4096 bytes of control data per message.

5. Local socket addresses

5.1. An AF_UNIX address is the C library’s struct sockaddr_un (sun_family, then a 108-byte sun_path). There are three kinds, told apart by addrlen and the first path byte:

Kind Form Name
Unnamed addrlen == sizeof(sa_family_t) none
Path name sun_path[0] != '\0' the bytes up to the first NUL or the end of addrlen
Abstract sun_path[0] == '\0', addrlen > sizeof(sa_family_t) exactly sun_path[1 .. addrlen - 2]; NUL bytes are part of it

5.2. Path names. bind(2) with a path name MUST create a new file of type S_IFSOCK with mode 0777 & ~umask, owned by the caller. It fails with EADDRINUSE if anything exists at the path, and with the usual path errors (ENOENT, ENOTDIR, EACCES for want of write and search permission on the directory, ENAMETOOLONG, ELOOP, EROFS). The file names the socket, not the reverse: closing the socket leaves the file, and removing the file leaves the socket working for already-connected peers. Renaming the file moves the name. connect(2) and sending to a path name require write permission on the file (EACCES); a socket file with no socket bound to it gives ECONNREFUSED; a file of another type gives ENOTSOCK. open(2) on a socket file fails with EOPNOTSUPP.

5.3. Abstract names (from Linux) exist only while a socket is bound to them. They are not files, carry no permissions, are not subject to chroot(2), and are released when the socket is closed. bind(2) to a name already bound fails with EADDRINUSE. Binding with an unnamed address (addrlen == sizeof(sa_family_t)) binds the socket to a fresh abstract name of five hexadecimal digits, as Linux does (“autobind”).

5.4. getsockname(2), getpeername(2), accept(2) and recvmsg(2) return: offsetof(sun_path) + strlen(path) + 1 for a path name, sizeof(sa_family_t) for an unnamed socket, and sizeof(sa_family_t) + 1 + length for an abstract name. A returned path name is the path as given to bind(2).

Rationale. The BSDs keep a sun_len byte and a 104-byte path; the C library here is musl, whose struct sockaddr_un is Linux’s, and the kernel follows the library. The abstract namespace is Linux’s; it is provided because ported software uses it and the planned Linux compatibility layer needs it.

6. Stream and sequenced-packet sockets

6.1. listen(2) marks a bound socket as accepting connections, with a queue of at most min(backlog, 128) connections not yet accepted; a negative or zero backlog means 1. listen(2) on an unbound socket fails with EINVAL.

6.2. connect(2) to a listening socket completes at once: the new connection enters the listener’s queue, and both sides may send immediately. It fails with ECONNREFUSED when the queue is full, as in the BSDs; it does not wait. accept(2) takes connections from the queue in arrival order.

6.3. Each direction of a connection has one buffer, sized by the receiver’s SO_RCVBUF (default 64 KiB, range 512 bytes to 1 MiB). A send blocks while the buffer lacks room.

6.4. SOCK_STREAM carries bytes without boundaries, except that a receive MUST NOT return data from two sends if the later one carried control data: control data stays attached to the bytes it was sent with.

6.5. SOCK_SEQPACKET keeps record boundaries. A record larger than the receive buffer given is truncated, the rest of it is discarded, and MSG_TRUNC is set in msg_flags. Every record ends with MSG_EOR.

6.6. Shutdown and close. After shutdown(SHUT_WR) the peer reads end-of-file once the buffer is drained. After shutdown(SHUT_RD) data arriving is discarded. When a side closes, the other reads the remaining data and then end-of-file, and its sends fail with EPIPE (§3.3.6). Connections still queued when a listener closes are reset: their reads fail with ECONNRESET.

7. Datagram sockets

7.1. Messages keep their boundaries. A message larger than the sender’s SO_SNDBUF (default 64 KiB) fails with EMSGSIZE. A message larger than the receiver’s buffer given is truncated and MSG_TRUNC is set.

7.2. Each receiving socket queues at most SO_RCVBUF bytes (default 64 KiB). A message that does not fit is dropped and the send fails with ENOBUFS, as in the BSDs.

7.3. connect(2) sets the default destination for sends without an address. It does not filter what the socket receives. Connecting to an unnamed address (AF_UNSPEC) removes the default destination.

7.4. When the default destination’s socket is closed, a send without an address fails with ENOTCONN; the sender may connect again.

7.5. A message from a sender with no name arrives with an unnamed address.

8. Descriptor passing

8.1. A control message of level SOL_SOCKET, type SCM_RIGHTS, carrying an array of int, sends the descriptions those descriptors refer to. Every descriptor MUST be open in the sender (EBADF otherwise). The kernel holds a reference to each description from the send until it is received or discarded; closing the sender’s descriptors does not affect a message in flight.

8.2. On receipt, each description is installed in the receiver at the lowest free descriptor, with FD_CLOEXEC set if MSG_CMSG_CLOEXEC was given. If the control buffer is too small for all of them, or the receiver cannot open more descriptors, the descriptions that were not installed are closed and MSG_CTRUNC is set.

8.3. A message discarded unread (by SHUT_RD, by closing its socket, or by a datagram overflow) closes the descriptions it carried.

8.4. Garbage collection. A socket descriptor in flight may make itself unreachable (sent over itself, or in a cycle of sockets). When a local socket that has descriptors in flight is closed, the kernel MUST run a mark-and-sweep collection over in-flight socket descriptions, as the BSDs’ unp_gc does, and close those no process can reach.

8.5. At most 1024 descriptions sent by one user may be in flight at once, and at most 4096 system-wide; a send beyond either limit fails with ETOOMANYREFS. Root is subject to the system-wide limit only.

9. Credentials

9.1. A process’s credentials here are its process ID, user ID and group ID. Until effective and saved IDs exist (SUDO.md), the effective IDs equal the real ones, and the group list is the single group ID.

9.2. Connection credentials are recorded when a connection is made: the connecting side records the listener’s credentials as of listen(2), the accepted side records the connector’s as of connect(2), and each end of a socketpair(2) records its creator’s. They are read by: 1. getpeereid(3) (all three BSDs): the peer’s effective user and group IDs; 2. getsockopt(fd, SOL_LOCAL, LOCAL_PEERCRED) (FreeBSD): a struct xucred (cr_version XUCRED_VERSION, cr_uid, cr_ngroups, cr_groups[16], cr_pid); 3. getsockopt(fd, SOL_SOCKET, SO_PEERCRED) (Linux): a struct ucred (pid, uid, gid).

Each fails with ENOTCONN on an unconnected socket and EINVAL on a socket of another family.

9.3. Per-message credentials, BSD form. A sender that includes a control message of type SCM_CREDS has its contents replaced by the kernel with the sender’s struct cmsgcred (cmcred_pid, cmcred_uid, cmcred_euid, cmcred_gid, cmcred_ngroups, cmcred_groups[16]), which the receiver gets unchanged. The sender cannot forge it.

9.4. Per-message credentials, Linux form. A receiver with SO_PASSCRED set receives an SCM_CREDENTIALS control message (struct ucred) with every message. A sender may supply one itself; the kernel MUST reject it with EPERM unless its pid is the sender’s and its uid and gid are the sender’s, or the sender is root. Without one, the kernel supplies the sender’s own.

9.5. Per-message credentials, receiver-requested (FreeBSD, NetBSD). A receiver that sets LOCAL_CREDS (level SOL_LOCAL) gets an SCM_CREDS control message holding the sender’s struct sockcred (sc_uid, sc_euid, sc_gid, sc_egid, sc_ngroups, sc_groups[]): with every message on a datagram socket, and with the first receive only on a stream or sequenced-packet socket. LOCAL_CREDS_PERSISTENT instead gives an SCM_CREDS2 control message holding struct sockcred2, which adds sc_version and sc_pid, with every message. The two options are mutually exclusive (EINVAL). While either is set, an SCM_CREDS message supplied by the sender (§9.3) is dropped, so the receiver sees one set of credentials, the kernel’s.

9.6. SOL_LOCAL, LOCAL_PEERCRED, LOCAL_CREDS, LOCAL_CREDS_PERSISTENT, SCM_CREDS, SCM_CREDS2, struct xucred, struct cmsgcred, struct sockcred, struct sockcred2 and getpeereid(3) are added to the C library. Their values MUST NOT collide with any SOL_*, SO_* or SCM_* value musl already defines.

10. Socket pairs

10.1. socketpair(AF_UNIX, type, 0, sv) accepts all three types and returns two unnamed sockets connected to each other. It replaces the pipe-based pair in sys/fs/pipe.rs, which is removed.

10.2. Any other domain fails with EOPNOTSUPP.

11. Internet sockets on the socket layer

11.1. UDP, TCP and raw ICMP become protocol-switch entries without changing their protocol behavior. They gain what the generic layer provides (§3.3): the MSG_* flags, timeouts, SO_ERROR, getpeername(2), honest address lengths.

11.2. An option or flag an AF_INET protocol does not implement fails with ENOPROTOOPT or EOPNOTSUPP.

12. Verification

12.1. tests/unix_syscall_smoke.rs with a C fixture using the musl API, one PASS/FAIL line per check, covering each numbered requirement of §§5–10 that can be observed from one boot: naming, permissions, stream/datagram/sequenced-packet semantics, shutdown, descriptor passing (including a descriptor sent over its own socket and collected), and every credential interface.

12.2. The existing udp, tcp, poll, ppoll, socketpair, ping and std networking tests MUST pass unchanged, and wget over HTTPS MUST still work.

13. Open questions

None.

Source: UNIX.md