Linux内核分析之进程间通信-03
This language version is unavailable; showing the other language.
29.1 POSIX 消息队列与 mqueue 文件系统
POSIX 消息队列把消息语义嫁接到 fd 世界:mq_open 返回的 mqd 可被 poll/epoll/close 处理,mq_notify 提供三件套中唯一的主动通知。实现上队列本体是 mqfs 内部文件系统的 inode,消息按优先级组织成红黑树,收发快路径以 pipelined 指针移交消除中间排队态。本节逐层拆解。
29.1.1 数据结构:mqfs inode 与优先级树
// ipc/mqueue.c:133-163
struct mqueue_inode_info {
spinlock_t lock;
struct inode vfs_inode; /* 就是 mqfs 的 inode */
wait_queue_head_t wait_q; /* poll 用等待队列 */
struct rb_root msg_tree; /* 按优先级组织的消息树 */
struct rb_node *msg_tree_rightmost; /* 最高优先级缓存 */
struct posix_msg_tree_node *node_cache;
struct mq_attr attr; /* mq_maxmsg/msgsize/currentmsgs */
struct sigevent notify; /* 异步通知注册 */
struct pid *notify_owner;
u32 notify_self_exec_id;
struct user_namespace *notify_user_ns;
struct ucounts *ucounts; /* 创建者记账 */
struct sock *notify_sock;
struct sk_buff *notify_cookie; /* netlink 通知载体 */
/* for tasks waiting for free space and messages, respectively */
struct ext_wait_queue e_wait_q[2]; /* [RECV] [SEND] 两条睡眠队列 */
unsigned long qsize; /* size of queue in memory (sum of all msgs) */
};
消息入树的规则:以 msg_prio 为键插入 msg_tree,同优先级按 FIFO 排在 posix_msg_tree_node 的链上;msg_tree_rightmost 缓存最大优先级节点,使"取最高优先级"从 O(log n) 树搜索降为一次指针解用——发送/摘取的更新路径同步维护该缓存。e_wait_q[2] 按接收者/发送者分队列睡眠(与 28.3 节 SysV 的三条具名链对照)。
队列本体是 mqfs 的 inode:每个 IPC namespace 在 mq_init_ns()(mqueue.c:465 注释"currently the only caller of mq_create_mount()")挂载一份内部 mqfs,/dev/mqueue/name 提供观测视图(cat 可见 attr,37 章伪文件系统家族),MQUEUE_I()(mqueue.c:166)从 inode 反查本体。ucounts 按创建用户记账(RLIMIT_MSGQUEUE),防止单用户以海量队列耗尽内核内存。
29.1.2 发送与接收:pipelined 快路径
// ipc/mqueue.c:1007-1025
/* pipelined_send() - send a message directly to the task waiting in
* sys_mq_timedreceive() (without inserting message into a queue).
*/
static inline void pipelined_send(struct wake_q_head *wake_q,
struct mqueue_inode_info *info,
struct msg_msg *message,
struct ext_wait_queue *receiver)
{
receiver->msg = message;
__pipelined_op(wake_q, info, receiver);
}
/* pipelined_receive() - if there is task waiting in sys_mq_timedsend()
* gets its message and put to the queue (we have one free place for sure). */
static inline void pipelined_receive(struct wake_q_head *wake_q,
struct mqueue_inode_info *info)
{
struct ext_wait_queue *sender = wq_get_first_waiter(info, SEND);
if (!sender) {
/* for poll */
wake_up_interruptible(&info->wait_q);
return;
}
if (msg_insert(sender->msg, info))
return;
__pipelined_op(wake_q, info, sender);
}
do_mq_timedsend 协议:
[1] fd 权限校验 (O_WRONLY/O_RDWR) + 消息 ≤ attr.mq_msgsize
[2] 持 info->lock:
有等待接收者 (e_wait_q[RECV])?
→ pipelined_send: receiver->msg = 消息,
wake_q 唤醒 — 消息不进树, 零中间排队态
否则队列未满: msg_insert 按优先级入树
队列满: IPC_NOWAIT → -EAGAIN; 否则挂 e_wait_q[SEND] 睡
[3] 接收者醒来自取 receiver->msg 拷回用户态
do_mq_timedreceive 对称:
树非空 → 摘"最高优先级+FIFO"消息
树空且有等待发送者 → pipelined_receive:
接收者的"腾出空位"直接转给队首发送者
(sender->msg 入树), 双方零无效唤醒
pipelined 协议的内存序配方与 28.3.3 节 SysV 的 MSG_BARRIER 注释同源(mqueue.c 的注释明确说明这是同一优化的起源地):移交指针以 smp_store_release 发布,接收方在 syscall 返回路径 READ_ONCE+smp_acquire__after_ctrl_dep 消费——醒后不拿锁直接确认结果,锁只在需要重睡或出队清理时才碰。wake_q(17.4 节批量唤醒结构)保证持锁期间不直接 wake(防唤醒的进程立即自旋抢锁)。
29.1.3 通知:mq_notify 的两种投递
// ipc/mqueue.c:1266(入口)
static int do_mq_notify(mqd_t mqdes, const struct sigevent *notification)
mq_notify 注册"消息到达时主动告知",sigevent 分三种投递:
SIGEV_SIGNAL: 消息到达 → send_sig(sv_signo, notify_owner)
(si_code = SI_MESGQ, si_value 携带 sv_value;
notify_owner/notify_user_ns 记录投递对象与
权限上下文 — 以注册者身份发信号)
SIGEV_THREAD: glibc 层模拟 — 内核只经 notify_sock/
notify_cookie 发 netlink 通知
(48 章 Netlink 的冷门消费者), glibc 的
辅助线程收到后调起用户回调函数
SIGEV_NONE: 仅注册语义, 配合 mq_getattr 轮询
一次性语义 (POSIX): 通知发出即注销 —
注册者必须"收到后重新 mq_notify" 才能继续收,
防止消息风暴期间的重复信号洪泛
remove_notification(mqueue.c:156 区声明)在 fd 关闭/队列销毁时清理注册——通知注册的绑定对象是 fd 而非进程,fd 关闭即失效,与 epoll 的生命周期纪律一致。
29.1.4 fd 化的红利与生命周期
epoll 集成是 mqueue 相对 SysV msg 的最大工程优势:mqueue_file_operations 的 poll 把 info->wait_q 织入事件循环(45.3 节),服务端线程可以同时监视消息队列、socket、pipe——事件驱动架构无需为 mqueue 单独拉线程。mq_timedreceive 的超时与 epoll 定时器构成双保险。
生命周期采用引用计数模型:fd 引用 + 队列非空双条件,最后关闭且排空者触发销毁——解决了 28.3.4 节 SysV"进程全退队列仍驻留"的泄漏顽疾。限额在 /proc/sys/fs/mqueue/{queues_max,msg_max,msgsize_max}(namespace 级),与 mq_open 的 attr 参数(mq_maxmsg/mq_msgsize,首个创建者定调)两级约束。
| 维度 | SysV msg (28.3) | POSIX mqueue (本节) |
|---|---|---|
| 命名 | key/id | /name 路径(mqfs) |
| 优先级 | 接收侧 msgtyp 过滤 | 发送侧 msg_prio 排序 |
| 通知 | 无 | mq_notify 信号/netlink + epoll |
| 持久性 | namespace 存活,进程退不删 | 引用计数(fd 全关+空即回收) |
| 事件集成 | 无 fd,不可 | poll/epoll/select 全支持 |
小结
POSIX 消息队列以 mqfs inode 为本体、优先级红黑树为消息组织(右端节点缓存加速最高优先级摘取)、双睡眠队列为阻塞支撑;pipelined_send/receive 把"有对端在等"的快路径压缩成指针移交(wake_q 批量唤醒 + release/acquire 免锁返回),mq_notify 的信号/netlink 双投递与一次性注销语义提供主动通知,fd 化使 mqueue 融入 epoll 生态。相较 SysV 版,它以命名规范化与引用计数生命周期换取了事件驱动集成能力——消息语义的两大实现至此对照完毕。