Linux内核分析之进程间通信-00
26.1 管道原理与 pipe_inode_info
管道的全部状态浓缩在一个 pipe_inode_info 里:一个互斥锁、两套等待队列、一个环形缓冲索引、读写者计数。它的精妙在于用最小状态承载完整的生产者-消费者语义——环形队列天然解决"谁先谁后",计数器天然解决"端的存在性",pipe_buf_operations 则让同一环可以装页缓存页、匿名页甚至通知消息。本节逐字段拆解。
26.1.1 pipe_buffer:环上的一个槽位
// include/linux/pipe_fs_i.h:26-33
struct pipe_buffer {
struct page *page; /* 数据所在页 (4KB) */
unsigned int offset, len; /* 页内偏移与有效长度 */
const struct pipe_buf_operations *ops; /* 该缓冲的家族 */
unsigned int flags;
unsigned long private; /* ops 私有数据 */
};
pipe_buffer 的三种典型身份:
[1] 匿名管道页: page = 伙伴系统匿名页
ops = anon_pipe_buf_ops, 全程内核态拷贝进出
[2] 页缓存引用: splice 从文件读入时, page 直接指向
页缓存页 (零拷贝!), ops = page_cache_pipe_buf_ops
->confirm 等待 IO, ->release = put_page
[3] 通知消息 (watch_queue): 装内核事件消息,
->release 走通知子系统
flags 的关键位 (pipe_fs_i.h:9-16):
PIPE_BUF_FLAG_CAN_MERGE 0x10 尾部可续写 (26.2.2 节)
PIPE_BUF_FLAG_PACKET 0x08 packet 模式 (读整包)
PIPE_BUF_FLAG_WHOLE 0x20 读必须整包或出错
PIPE_BUF_FLAG_GIFT 0x04 vmsplice 赠与页
ops 指针的多态是管道支持零拷贝的机关:读写路径只认 pipe_buf_confirm/release 接口,不关心页背后是匿名内存还是页缓存——splice() 的实现(36 章 BIO 视角)正是把文件页"挂"进环、再让对端消费,全程无数据拷贝。
26.1.2 环形索引:head/tail 的不变式
// include/linux/pipe_fs_i.h:45-53(head_tail 合并字注释要点)
union pipe_index {
unsigned long head_tail; /* 合并成一个字 */
struct {
pipe_index_t head; /* 生产点 */
pipe_index_t tail; /* 消费点 */
};
};
// include/linux/pipe_fs_i.h:84-125(节选)
struct pipe_inode_info {
struct mutex mutex; /* 串行化读写者 */
wait_queue_head_t rd_wait, wr_wait; /* 读者/写者睡眠队列 */
union pipe_index; /* head/tail 内联展开 */
unsigned int max_usage; /* 历史峰值(账目用) */
unsigned int ring_size; /* 槽位总数 (默认 16) */
unsigned int nr_accounted; /* 计入 RLIMIT 的槽位数 */
unsigned int readers; /* 在读端数 */
unsigned int writers; /* 在写端数 */
unsigned int files; /* 引用此管道的 file 数 */
unsigned int r_counter; /* 读者累计计数(开闭配对) */
unsigned int w_counter; /* 写者累计计数 */
bool poll_usage; /* 有 poller 时限制 resize */
struct page *tmp_page[2]; /* 免锁的临时页缓存 */
struct fasync_struct *fasync_readers; /* O_ASYNC 通知链 */
struct fasync_struct *fasync_writers;
struct pipe_buffer *bufs; /* 环形数组本体 */
struct user_struct *user; /* 记账主体 */
...
};
环形不变式只有三条,全部读写路径围绕它们展开:
不变式 (head/tail 单调递增, 以 ring_size 取模定位):
pipe_empty(head, tail): head == tail → 读阻塞/写就绪
pipe_full(head, tail, ring_size):
head - tail == ring_size → 写阻塞/读就绪
定位: pipe_buf(pipe, idx) = &bufs[idx & (ring_size-1)]
(ring_size 恒为 2 的幂 → 取模即与运算)
head_tail 合并成单字的用意: 26.2 节的读路径用
smp_load_acquire 读 head —— 单字装载保证
"读到 head 必同时读到合法 tail" 的原子视图,
这是 15.2 节 acquire/release 语义的直接应用
readers/writers 与 r_counter/w_counter 的分工:前者是"当前在场的端数"(决定 EPIPE/EOF),后者是"历史上打开过的累计数"——FIFO 的 fifo_open 用 counter 差值实现"等待对端出现"的配对语义(26.3 节)。files 计数管 file 对象生命周期,poll_usage 在有 poller 时禁止 pipe_resize_ring 换环(poller 缓存的就绪状态会失配)。
26.1.3 创建:从 pipe2 到两个 fd
// fs/pipe.c:926(要点)
int create_pipe_files(struct file **res, int flags)
{ ...
inode = ... /* 新建 pipefs inode (内部文件系统, 37.4 节) */
...
f = alloc_file_pseudo(inode, pipe_mnt, "", O_WRONLY | flags,
&pipeanon_fops);
...
f->private_data = inode->i_pipe; /* 两个 fd 共享同一 pipe */
res[0] = alloc_file_pseudo(..., O_RDONLY | flags, &pipeanon_fops);
...
}
// fs/pipe.c:1032-1061
static int do_pipe2(int __user *fildes, int flags)
{ ... /* 分配两个 fd, fildes[0]=读 fildes[1]=写, 有序回滚 */ }
SYSCALL_DEFINE2(pipe2, int __user *, fildes, int, flags) /* :1054 */
{ return do_pipe2(fildes, flags); }
SYSCALL_DEFINE1(pipe, int __user *, fildes) /* :1061 */
{ return do_pipe2(fildes, 0); }
管道的 inode 居住在 pipefs——一个没有挂载点的内部文件系统(37.4 节伪文件系统家族),d_path 里显示为 pipe:[inode号]。fd 的继承由 fork 的文件表复制免费获得——shell 建管道、fork、exec 三步完成流水线,这也是管道天然面向"有亲缘进程"的原因,FIFO(26.3 节)补上名字的缺口。
alloc_pipe_info()(pipe.c)在创建时按 PIPE_DEF_BUFFERS=16(pipe_fs_i.h:5)分配环,并受 /proc/sys/fs/pipe-max-size(默认 1MB)与 pipe-user-pages-soft(每用户总页数配额,防内存耗尽攻击)双重限额——user/nr_accounted 字段即为此记账。
小结
管道 = 环形页缓冲 + 两套等待队列 + 端计数。pipe_buffer 经 ops 指针实现"匿名页/页缓存/通知消息"三种身份的多态,head/tail 合并字支撑 acquire 原子视图,readers/writers 与双 counter 分别承担存在性与开闭配对语义。pipefs 内部文件系统提供 inode,fork 免费传播 fd——最小状态承载完整生产者-消费者协议。下一节沿读写路径验证这套不变式如何被执行。
26.2 管道读写与缓冲区管理
读写路径是管道协议的执行现场:anon_pipe_read() 从 tail 消费、anon_pipe_write() 向 head 生产,两端在 pipe->mutex 串行化下依"空/满/对端存在"三态决定阻塞、唤醒或报错。本节沿两条主路径逐段分析,并覆盖尾部合并、SIGPIPE、poll 就绪判定与 F_SETPIPE_SZ 动态扩容。
26.2.1 读路径:anon_pipe_read
// fs/pipe.c:269-405(节选)
anon_pipe_read(struct kiocb *iocb, struct iov_iter *to)
{
size_t total_len = iov_iter_count(to);
struct pipe_inode_info *pipe = filp->private_data;
bool wake_writer = false, wake_next_reader = false;
ssize_t ret;
/* Null read succeeds. */
if (unlikely(total_len == 0))
return 0;
ret = 0;
mutex_lock(&pipe->mutex);
/*
* We only wake up writers if the pipe was full when we started reading
* and it is no longer full after reading to avoid unnecessary wakeups.
* But when we do wake up writers, we do so using a sync wakeup
* (WF_SYNC), because we want them to get going and generate more data.
*/
for (;;) {
/* Read ->head with a barrier vs post_one_notification() */
unsigned int head = smp_load_acquire(&pipe->head); /* :296 */
unsigned int tail = pipe->tail;
...
if (!pipe_empty(head, tail)) {
struct pipe_buffer *buf = pipe_buf(pipe, tail); /* :312 */
size_t chars = buf->len;
...
if (chars > total_len) {
if (buf->flags & PIPE_BUF_FLAG_WHOLE) {
if (ret == 0)
ret = -ENOBUFS;
break; /* packet 模式: 整包或错 */
}
chars = total_len; /* 普通模式: 部分消费 */
}
...
written = copy_page_to_iter(buf->page, buf->offset + buf->tail...);
ret += chars;
total_len -= chars;
buf->offset += chars;
buf->len -= chars;
...
if (!buf->len) {
pipe_buf_release(pipe, buf); /* 空槽归还 */
tail++;
pipe->tail = tail;
}
if (!total_len)
break;
continue; /* 迭代继续消费下一个槽 */
}
... /* 空环处理, 见下 */
}
...
}
读循环的消费协议:smp_load_acquire 取 head(:296 注释"Read ->head with a barrier")→ 非空则从 tail 槽 copy_page_to_iter 拷出 → 部分消费时只动 offset/len(一个槽可被多次 read 撕开读),整槽耗尽才 pipe_buf_release 归还页并推进 tail。PIPE_BUF_FLAG_WHOLE(packet 模式,fcntl(F_SETPIPE_SZ) 关联的 O_DIRECT 式语义)禁止拆包——要么整包读走要么 -ENOBUFS。
空环的三态出口(循环的后半段,:340 区):
空环时读者怎么办?
readers == 0 ? 不可能 (自己就是读者)
writers == 0 → EOF: 返回已读字节数; 一字节没读过则 0
(管道关闭的经典信号 — shell 流水线结束)
O_NONBLOCK → 返回 -EAGAIN
其余 → prepare_to_wait(rd_wait) 睡眠,
醒来重进循环复查 (惊群安全: 唤醒后
需重新拿 mutex, 状态可能已被别的读者吃掉)
26.2.2 写路径:anon_pipe_write 与尾部合并
// fs/pipe.c:431-602(节选)
anon_pipe_write(struct kiocb *iocb, struct iov_iter *from)
{
...
mutex_lock(&pipe->mutex);
if (!pipe->readers) { /* :457 */
if ((iocb->ki_flags & IOCB_NOSIGNAL) == 0)
send_sig(SIGPIPE, current, 0); /* 经典之死 */
ret = -EPIPE;
goto out;
}
/*
* If it wasn't empty we try to merge new data into
* the last buffer.
* That naturally merges small writes, but it also
* page-aligns the rest of the writes for large writes.
*/
head = pipe->head;
was_empty = pipe_empty(head, pipe->tail);
chars = total_len & (PAGE_SIZE-1); /* :477 尾数 */
if (chars && !was_empty) {
struct pipe_buffer *buf = pipe_buf(pipe, head - 1);
int offset = buf->offset + buf->len;
if ((buf->flags & PIPE_BUF_FLAG_CAN_MERGE) &&
offset + chars <= PAGE_SIZE) { /* :484 */
ret = pipe_buf_confirm(pipe, buf);
...
ret = copy_page_from_iter(buf->page, offset, chars, from);
...
buf->len += ret;
if (!iov_iter_count(from))
goto out; /* 全部塞进尾槽, 完事 */
}
}
for (;;) {
...
head = pipe->head;
if (!pipe_full(head, pipe->tail, pipe->max_usage)) {
page = anon_pipe_get_page(pipe); /* 新页供货 */
copied = copy_page_from_iter(page, 0, PAGE_SIZE, from);
...
pipe->head = head + 1; /* 发布新槽 */
...
continue;
}
... /* 满环: EAGAIN 或 睡 wr_wait */
}
}
三个写侧策略:
- SIGPIPE 协议(:457-461):写端发现读者绝迹,先
send_sig(SIGPIPE)再返-EPIPE——SIGPIPE默认终止进程,shell 流水线靠它让上游死得体面;IOCB_NOSIGNAL(或进程忽略 SIGPIPE)时只返回错误码。 - 尾部合并(:477-497):
total_len & (PAGE_SIZE-1)算出"尾数"(最后一页不满的部分),若尾槽带CAN_MERGE且放得下就原地续写——高频小写的缓冲友好路径,避免一次 write 烧掉多个整页槽。anon_pipe_write建新槽时给最后一段打上CAN_MERGE,后续槽不带(只有尾槽可续)。 - 满环出口:非阻塞返
-EAGAIN;阻塞睡眠于wr_wait,被读者唤醒后复查pipe->readers(循环顶再次出现 :505 的读者绝迹检查——睡眠期间对端可能关闭,醒来第一件事是重新评估生死)。
写入原子性:write(fd, buf, n) 中 n ≤ PIPE_BUF(4KB,POSIX 保证)时在阻塞管道上原子——实现依据是写者持有 pipe->mutex 直到"全部数据入环或睡眠后完成",读者不可能看到半条 ≤4KB 的消息交错。
26.2.3 poll 与就绪判定
// fs/pipe.c:660(要点)
pipe_poll(struct file *filp, poll_table *wait)
{
...
poll_wait(filp, &pipe->rd_wait, wait); /* 两队都注册 */
poll_wait(filp, &pipe->wr_wait, wait);
...
head = smp_load_acquire(&pipe->head);
tail = smp_load_acquire(&pipe->tail);
mask = 0;
if (filp->f_mode & FMODE_READ) {
if (!pipe_empty(head, tail))
mask |= EPOLLIN | EPOLLRDNORM; /* 有数据: 可读 */
if (!pipe->writers && filp->f_pipe != pipe->w_counter)
mask |= EPOLLHUP; /* 写端全关: 挂断 */
...
}
if (filp->f_mode & FMODE_WRITE) {
if (!pipe_full(head, tail, ring_size))
mask |= EPOLLOUT | EPOLLWRNORM; /* 有空位: 可写 */
if (!pipe->readers)
mask |= EPOLLERR; /* 读端全关: 错误 */
}
return mask;
}
poll 的事件表与读写的三态严格对应:EPOLLIN ⇔ 非空、EPOLLOUT ⇔ 非满、写端绝迹 → EPOLLHUP、读端绝迹 → EPOLLERR。epoll(45.3 节)把这两条等待队列织进事件循环——管道是事件驱动服务进程的标准管道(字面意义),网络代理把 socket 数据经 splice 灌进 pipe 环再 epoll 等待就是这个模型。
26.2.4 动态扩容:pipe_resize_ring
// fs/pipe.c:1291(签名)
int pipe_resize_ring(struct pipe_inode_info *pipe, unsigned int nr_slots)
fcntl(fd, F_SETPIPE_SZ, size)(fs/pipe.c:1439)把环重排到 nr_slots = size/PAGE_SIZE(上限 pipe-max-size 或 CAP_SYS_RESOURCE):分配新数组 → 把旧环从 tail 起线性搬运(环形→线性→环形的两段拷贝)→ 换指针。约束链:新容量 ≥ 当前未消费数据量、无 poller(poll_usage 检查)、与 pipe-user-pages 配额协商。缩小也允许(数据仍装得下时)——这是 26.1.3 节默认 16 槽之外,高吞吐场景(如日志中继)调到 64/256 槽的入口;反向,F_GETPIPE_SZ(:1442)读取,FIONREAD(:625)报"还能读多少字节"。
小结
读路径以 smp_load_acquire 取视图、从 tail 逐槽消费、空环按"EOF/EAGAIN/睡眠"三态出口;写路径以 SIGPIPE 守灵、尾槽 CAN_MERGE 吸收小写、满环睡眠并在醒来后重判生死;≤4KB 写的原子性由 mutex 持有跨度保证。poll 把三态映射成 EPOLL 事件,F_SETPIPE_SZ 经 pipe_resize_ring 让环在 16 槽之外弹性伸缩。管道协议的全部机关至此完整;下一节看 FIFO 如何把同一设施接到文件名上。
26.3 FIFO (命名管道)
FIFO 与匿名管道共享同一套 pipe_inode_info 设施与读写代码——差别只在"端从哪来":匿名管道的两个 fd 由 pipe2() 一次发亲,FIFO 的两端则由任意进程各自 open() 同一个路径名获得。本节分析 mknod 建立的 inode 如何嫁接 pipe 操作、fifo_open() 的 POSIX 打开四象限,以及读写行为与匿名管道的细微差别。
26.3.1 嫁接:从 inode 到 pipe_inode_info
// fs/pipe.c:1121-1160(节选)
static int fifo_open(struct inode *inode, struct file *filp)
{
bool is_pipe = inode->i_fop == &pipeanon_fops;
struct pipe_inode_info *pipe;
int ret;
filp->f_pipe = 0;
spin_lock(&inode->i_lock);
if (inode->i_pipe) { /* 已有管道: 附着 */
pipe = inode->i_pipe;
pipe->files++;
spin_unlock(&inode->i_lock);
} else {
spin_unlock(&inode->i_lock);
pipe = alloc_pipe_info(); /* 第一个打开者创建 */
if (!pipe)
return -ENOMEM;
pipe->files = 1;
spin_lock(&inode->i_lock);
if (unlikely(inode->i_pipe)) {
/* 双开竞态: 用对方刚建的, 丢弃自己的 */
inode->i_pipe->files++;
spin_unlock(&inode->i_lock);
free_pipe_info(pipe);
pipe = inode->i_pipe;
} else {
inode->i_pipe = pipe;
spin_unlock(&inode->i_lock);
}
}
filp->private_data = pipe;
/* OK, we have a pipe and it's pinned down */
mutex_lock(&pipe->mutex);
/* We can only do regular read/write on fifos */
stream_open(inode, filp);
...
}
FIFO 的 inode 由 mknod(S_IFIFO) 创建,i_fop 指向 pipefifo_fops(pipe.c:1258 的 .open = fifo_open)。第一个打开者创建 pipe 实体、后续打开者附着——inode->i_pipe 指针 + i_lock 自旋锁的"先查后建再复核"模式是内核单实例创建的标准写法(与 21 章缓存创建同构)。管道实体的生命周期随 inode(路径存在即存活),而非随 fd——这是与匿名管道(最后一个 fd 关闭即亡)的根本差异。
26.3.2 打开四象限:POSIX 阻塞规则
// fs/pipe.c:1167-1245(四象限判定, 节选)
switch (filp->f_mode & (FMODE_READ | FMODE_WRITE)) {
case FMODE_READ:
/*
* O_RDONLY
* POSIX.1 says that O_NONBLOCK means return with the FIFO
* opened, even when there is no process writing the FIFO.
*/
pipe->r_counter++;
if (pipe->readers++ == 0)
wake_up_partner(pipe);
if (!is_pipe && !pipe->writers) { /* 还没有写者 */
if ((filp->f_flags & O_NONBLOCK)) {
/* suppress EPOLLHUP until we have
* seen a writer */
filp->f_pipe = pipe->w_counter;
} else {
if (wait_for_partner(pipe, &pipe->w_counter))
goto err_rd; /* 阻塞等写者 */
}
}
break;
case FMODE_WRITE:
/*
* O_WRONLY
* POSIX.1 says that O_NONBLOCK means return -1 with
* errno=ENXIO when there is no process reading the FIFO.
*/
ret = -ENXIO;
if (!is_pipe && (filp->f_flags & O_NONBLOCK) && !pipe->readers)
goto err; /* 非阻塞且无读者: ENXIO */
pipe->w_counter++;
if (!pipe->writers++)
wake_up_partner(pipe);
if (!is_pipe && !pipe->readers) {
if (wait_for_partner(pipe, &pipe->r_counter))
goto err_wr; /* 阻塞等读者 */
}
break;
case FMODE_READ | FMODE_WRITE:
/*
* O_RDWR
* POSIX.1 leaves this case "undefined" when O_NONBLOCK is set.
* This implementation will NEVER block on a O_RDWR open,
* since the process can at least talk to itself.
*/
pipe->readers++;
pipe->writers++;
wake_up_partner(pipe);
break;
}
...
四象限语义表(is_pipe=false 即 FIFO 场景):
| 打开方式 | 对端已存在 | 对端不存在 |
|---|---|---|
O_RDONLY 阻塞 |
立即成功 | 阻塞直到写者 open |
O_RDONLY+O_NONBLOCK |
立即成功 | 立即成功(:1188-1191 以 f_pipe 压住 EPOLLHUP 直到真见写者) |
O_WRONLY 阻塞 |
立即成功 | 阻塞直到读者 open |
O_WRONLY+O_NONBLOCK |
立即成功 | 失败 -ENXIO(:1204-1206——不对称的根源:写没有对端毫无意义) |
O_RDWR |
立即成功 | 立即成功(自己就是对端,:1220-1229 注释自证) |
wait_for_partner 的配对协议用 r_counter/w_counter 实现:打开者睡眠,对端 open 时 wake_up_partner 并推进 counter;醒来比对 counter——计数变了即"对端来过"。这解决了"先到者如何知道后到者已到"的握手问题,与 17.4 节 completion 是同一思想的手工实现。
O_RDONLY|O_NONBLOCK 的 f_pipe = pipe->w_counter 技巧(:1190)值得玩味:此刻 w_counter 是"上次写者关闭后的计数",poll 里 filp->f_pipe != pipe->w_counter 的比较使得"从未见过写者"的读端不报 HUP(否则空 FIFO 读端会立刻收到 EPOLLHUP,epoll 应用无法部署"等服务起来再写"的模式)。
26.3.3 FIFO 的读写差异
// fs/pipe.c:407-409
fifo_pipe_read(struct kiocb *iocb, struct iov_iter *to)
{
int ret = anon_pipe_read(iocb, to);
...
}
// fs/pipe.c:604-606
fifo_pipe_write(struct kiocb *iocb, struct iov_iter *from)
{
int ret = anon_pipe_write(iocb, from);
...
}
读写本身是匿名管道代码的薄包装(fifo_pipe_read/write 仅处理 FIFO 特有的状态修正),行为差异集中在边界:
| 差异点 | 匿名管道 | FIFO |
|---|---|---|
| 实体生命周期 | 最后 fd 关闭即亡 | 随路径 inode 存活(rm 才消亡) |
| EOF 语义 | 写端 fd 全关 → 读到 0 | 写端 open 的进程全退出 → 读到 0;"没人打开写端"与"写完了"由打开阻塞规则区分 |
| 多进程同时 open 写端 | 天然共享 | 同左(writers 计数) |
| 持久性 | 无 | 文件系统可见(ls -l 显示 p 类型) |
"FIFO 卡死"的经典事故正源于打开阻塞:服务端 open(fifo, O_RDONLY) 先启动、客户端晚启动,服务端阻塞在 wait_for_partner 看似"假死";正确姿势是服务端 O_RDONLY|O_NONBLOCK 打开后立即 epoll 注册(26.3.2 节的 f_pipe 技巧保证不误报 HUP)——这也是 systemd socket-activation 与多数守护进程的标准写法。
26.3.4 实用模式速查
[1] 单工日志通道:
应用 → O_WRONLY 打开 /run/log/pipe ← 日志聚合器 O_RDONLY
(聚合器先起: O_RDONLY|O_NONBLOCK + epoll)
[2] 双向通信: 两条 FIFO (fifo_req / fifo_resp)
FIFO 本身是单工的 — 环形队列没有"回话"概念
[3] 与匿名管道的选择:
有亲缘 → pipe2 (无路径管理成本)
无亲缘+一次性 → FIFO
无亲缘+消息边界/优先级 → 换 POSIX 消息队列 (29 章)
[4] 容量调优 (26.2.4 节):
高吞吐中继: fcntl(F_SETPIPE_SZ, 1<<20) 扩到 256 槽
内存受限: 缩小或依赖默认 64KB
小结
FIFO 把管道嫁接到文件系统:fifo_open 以"首开者建实体、后开者附着"的单实例模式挂起 pipe_inode_info,以 r/w counter 配对协议实现 POSIX 打开四象限(读端非阻塞不报错、写端非阻塞报 ENXIO 的不对称、O_RDWR 自圆其说),f_pipe 技巧让非阻塞读端在等写者期间不误报 EPOLLHUP。读写路径完全复用匿名管道代码,EOF 语义因"打开者计数"而更精细。第五部分的通信原语从最简单的管道开始;下一章进入信号——另一种语义完全不同的异步通知机制。