Linux内核分析之进程间通信-03

29.1 POSIX 消息队列与 mqueue 文件系统

POSIX 消息队列把消息语义嫁接到 fd 世界:mq_open 返回的 mqd 可被 poll/epoll/close 处理,mq_notify 提供三件套中唯一的主动通知。实现上队列本体是 mqfs 内部文件系统的 inode,消息按优先级组织成红黑树,收发快路径以 pipelined 指针移交消除中间排队态。本节逐层拆解。


29.1.1 数据结构:mqfs inode 与优先级树

// ipc/mqueue.c:133-163
struct mqueue_inode_info {
    spinlock_t lock;
    struct inode vfs_inode;         /* 就是 mqfs 的 inode */
    wait_queue_head_t wait_q;       /* poll 用等待队列 */

    struct rb_root msg_tree;        /* 按优先级组织的消息树 */
    struct rb_node *msg_tree_rightmost; /* 最高优先级缓存 */
    struct posix_msg_tree_node *node_cache;
    struct mq_attr attr;            /* mq_maxmsg/msgsize/currentmsgs */

    struct sigevent notify;         /* 异步通知注册 */
    struct pid *notify_owner;
    u32 notify_self_exec_id;
    struct user_namespace *notify_user_ns;
    struct ucounts *ucounts;        /* 创建者记账 */
    struct sock *notify_sock;
    struct sk_buff *notify_cookie;      /* netlink 通知载体 */

    /* for tasks waiting for free space and messages, respectively */
    struct ext_wait_queue e_wait_q[2];  /* [RECV] [SEND] 两条睡眠队列 */

    unsigned long qsize;    /* size of queue in memory (sum of all msgs) */
};

消息入树的规则:以 msg_prio 为键插入 msg_tree,同优先级按 FIFO 排在 posix_msg_tree_node 的链上;msg_tree_rightmost 缓存最大优先级节点,使"取最高优先级"从 O(log n) 树搜索降为一次指针解用——发送/摘取的更新路径同步维护该缓存。e_wait_q[2] 按接收者/发送者分队列睡眠(与 28.3 节 SysV 的三条具名链对照)。

队列本体是 mqfs 的 inode:每个 IPC namespace 在 mq_init_ns()(mqueue.c:465 注释"currently the only caller of mq_create_mount()")挂载一份内部 mqfs,/dev/mqueue/name 提供观测视图(cat 可见 attr,37 章伪文件系统家族),MQUEUE_I()(mqueue.c:166)从 inode 反查本体。ucounts 按创建用户记账(RLIMIT_MSGQUEUE),防止单用户以海量队列耗尽内核内存。

29.1.2 发送与接收:pipelined 快路径

// ipc/mqueue.c:1007-1025
/* pipelined_send() - send a message directly to the task waiting in
 * sys_mq_timedreceive() (without inserting message into a queue).
 */
static inline void pipelined_send(struct wake_q_head *wake_q,
                  struct mqueue_inode_info *info,
                  struct msg_msg *message,
                  struct ext_wait_queue *receiver)
{
    receiver->msg = message;
    __pipelined_op(wake_q, info, receiver);
}

/* pipelined_receive() - if there is task waiting in sys_mq_timedsend()
 * gets its message and put to the queue (we have one free place for sure). */
static inline void pipelined_receive(struct wake_q_head *wake_q,
                     struct mqueue_inode_info *info)
{
    struct ext_wait_queue *sender = wq_get_first_waiter(info, SEND);

    if (!sender) {
        /* for poll */
        wake_up_interruptible(&info->wait_q);
        return;
    }
    if (msg_insert(sender->msg, info))
        return;

    __pipelined_op(wake_q, info, sender);
}
do_mq_timedsend 协议:

 [1] fd 权限校验 (O_WRONLY/O_RDWR) + 消息 ≤ attr.mq_msgsize
 [2] 持 info->lock:
     有等待接收者 (e_wait_q[RECV])?
       → pipelined_send: receiver->msg = 消息,
         wake_q 唤醒 — 消息不进树, 零中间排队态
     否则队列未满: msg_insert 按优先级入树
     队列满: IPC_NOWAIT → -EAGAIN; 否则挂 e_wait_q[SEND] 睡
 [3] 接收者醒来自取 receiver->msg 拷回用户态

 do_mq_timedreceive 对称:
     树非空 → 摘"最高优先级+FIFO"消息
     树空且有等待发送者 → pipelined_receive:
        接收者的"腾出空位"直接转给队首发送者
        (sender->msg 入树), 双方零无效唤醒

pipelined 协议的内存序配方与 28.3.3 节 SysV 的 MSG_BARRIER 注释同源(mqueue.c 的注释明确说明这是同一优化的起源地):移交指针以 smp_store_release 发布,接收方在 syscall 返回路径 READ_ONCE+smp_acquire__after_ctrl_dep 消费——醒后不拿锁直接确认结果,锁只在需要重睡或出队清理时才碰。wake_q(17.4 节批量唤醒结构)保证持锁期间不直接 wake(防唤醒的进程立即自旋抢锁)。

29.1.3 通知:mq_notify 的两种投递

// ipc/mqueue.c:1266(入口)
static int do_mq_notify(mqd_t mqdes, const struct sigevent *notification)

mq_notify 注册"消息到达时主动告知",sigevent 分三种投递:

 SIGEV_SIGNAL: 消息到达 → send_sig(sv_signo, notify_owner)
              (si_code = SI_MESGQ, si_value 携带 sv_value;
               notify_owner/notify_user_ns 记录投递对象与
               权限上下文 — 以注册者身份发信号)

 SIGEV_THREAD: glibc 层模拟 — 内核只经 notify_sock/
              notify_cookie 发 netlink 通知
              (48 章 Netlink 的冷门消费者), glibc 的
              辅助线程收到后调起用户回调函数

 SIGEV_NONE:  仅注册语义, 配合 mq_getattr 轮询

 一次性语义 (POSIX): 通知发出即注销 —
 注册者必须"收到后重新 mq_notify" 才能继续收,
 防止消息风暴期间的重复信号洪泛

remove_notification(mqueue.c:156 区声明)在 fd 关闭/队列销毁时清理注册——通知注册的绑定对象是 fd 而非进程,fd 关闭即失效,与 epoll 的生命周期纪律一致。

29.1.4 fd 化的红利与生命周期

epoll 集成是 mqueue 相对 SysV msg 的最大工程优势:mqueue_file_operations 的 poll 把 info->wait_q 织入事件循环(45.3 节),服务端线程可以同时监视消息队列、socket、pipe——事件驱动架构无需为 mqueue 单独拉线程。mq_timedreceive 的超时与 epoll 定时器构成双保险。

生命周期采用引用计数模型:fd 引用 + 队列非空双条件,最后关闭且排空者触发销毁——解决了 28.3.4 节 SysV"进程全退队列仍驻留"的泄漏顽疾。限额在 /proc/sys/fs/mqueue/{queues_max,msg_max,msgsize_max}(namespace 级),与 mq_open 的 attr 参数(mq_maxmsg/mq_msgsize,首个创建者定调)两级约束。

维度 SysV msg (28.3) POSIX mqueue (本节)
命名 key/id /name 路径(mqfs)
优先级 接收侧 msgtyp 过滤 发送侧 msg_prio 排序
通知 无 mq_notify 信号/netlink + epoll
持久性 namespace 存活,进程退不删 引用计数(fd 全关+空即回收)
事件集成 无 fd,不可 poll/epoll/select 全支持

小结

POSIX 消息队列以 mqfs inode 为本体、优先级红黑树为消息组织(右端节点缓存加速最高优先级摘取)、双睡眠队列为阻塞支撑;pipelined_send/receive 把"有对端在等"的快路径压缩成指针移交(wake_q 批量唤醒 + release/acquire 免锁返回),mq_notify 的信号/netlink 双投递与一次性注销语义提供主动通知,fd 化使 mqueue 融入 epoll 生态。相较 SysV 版,它以命名规范化与引用计数生命周期换取了事件驱动集成能力——消息语义的两大实现至此对照完毕。