mirror of
https://github.com/linux-msm/laptops-kernel.git
synced 2026-08-13 14:19:53 -07:00
Merge tag 'vfs-6.17-rc1.pidfs' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull pidfs updates from Christian Brauner:
- persistent info
Persist exit and coredump information independent of whether anyone
currently holds a pidfd for the struct pid.
The current scheme allocated pidfs dentries on-demand repeatedly.
This scheme is reaching it's limits as it makes it impossible to pin
information that needs to be available after the task has exited or
coredumped and that should not be lost simply because the pidfd got
closed temporarily. The next opener should still see the stashed
information.
This is also a prerequisite for supporting extended attributes on
pidfds to allow attaching meta information to them.
If someone opens a pidfd for a struct pid a pidfs dentry is allocated
and stashed in pid->stashed. Once the last pidfd for the struct pid
is closed the pidfs dentry is released and removed from pid->stashed.
So if 10 callers create a pidfs dentry for the same struct pid
sequentially, i.e., each closing the pidfd before the other creates a
new one then a new pidfs dentry is allocated every time.
Because multiple tasks acquiring and releasing a pidfd for the same
struct pid can race with each another a task may still find a valid
pidfs entry from the previous task in pid->stashed and reuse it. Or
it might find a dead dentry in there and fail to reuse it and so
stashes a new pidfs dentry. Multiple tasks may race to stash a new
pidfs dentry but only one will succeed, the other ones will put their
dentry.
The current scheme aims to ensure that a pidfs dentry for a struct
pid can only be created if the task is still alive or if a pidfs
dentry already existed before the task was reaped and so exit
information has been was stashed in the pidfs inode.
That's great except that it's buggy. If a pidfs dentry is stashed in
pid->stashed after pidfs_exit() but before __unhash_process() is
called we will return a pidfd for a reaped task without exit
information being available.
The pidfds_pid_valid() check does not guard against this race as it
doens't sync at all with pidfs_exit(). The pid_has_task() check might
be successful simply because we're before __unhash_process() but
after pidfs_exit().
Introduce a new scheme where the lifetime of information associated
with a pidfs entry (coredump and exit information) isn't bound to the
lifetime of the pidfs inode but the struct pid itself.
The first time a pidfs dentry is allocated for a struct pid a struct
pidfs_attr will be allocated which will be used to store exit and
coredump information.
If all pidfs for the pidfs dentry are closed the dentry and inode can
be cleaned up but the struct pidfs_attr will stick until the struct
pid itself is freed. This will ensure minimal memory usage while
persisting relevant information.
The new scheme has various advantages. First, it allows to close the
race where we end up handing out a pidfd for a reaped task for which
no exit information is available. Second, it minimizes memory usage.
Third, it allows to remove complex lifetime tracking via dentries
when registering a struct pid with pidfs. There's no need to get or
put a reference. Instead, the lifetime of exit and coredump
information associated with a struct pid is bound to the lifetime of
struct pid itself.
- extended attributes
Now that we have a way to persist information for pidfs dentries we
can start supporting extended attributes on pidfds. This will allow
userspace to attach meta information to tasks.
One natural extension would be to introduce a custom pidfs.* extended
attribute space and allow for the inheritance of extended attributes
across fork() and exec().
The first simple scheme will allow privileged userspace to set
trusted extended attributes on pidfs inodes.
- Allow autonomous pidfs file handles
Various filesystems such as pidfs and drm support opening file
handles without having to require a file descriptor to identify the
filesystem. The filesystem are global single instances and can be
trivially identified solely on the information encoded in the file
handle.
This makes it possible to not have to keep or acquire a sentinal file
descriptor just to pass it to open_by_handle_at() to identify the
filesystem. That's especially useful when such sentinel file
descriptor cannot or should not be acquired.
For pidfs this means a file handle can function as full replacement
for storing a pid in a file. Instead a file handle can be stored and
reopened purely based on the file handle.
Such autonomous file handles can be opened with or without specifying
a a file descriptor. If no proper file descriptor is used the
FD_PIDFS_ROOT sentinel must be passed. This allows us to define
further special negative fd sentinels in the future.
Userspace can trivially test for support by trying to open the file
handle with an invalid file descriptor.
- Allow pidfds for reaped tasks with SCM_PIDFD messages
This is a logical continuation of the earlier work to create pidfds
for reaped tasks through the SO_PEERPIDFD socket option merged in
923ea4d448 ("Merge patch series "net, pidfs: enable handing out
pidfds for reaped sk->sk_peer_pid"").
- Two minor fixes:
* Fold fs_struct->{lock,seq} into a seqlock
* Don't bother with path_{get,put}() in unix_open_file()
* tag 'vfs-6.17-rc1.pidfs' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (37 commits)
don't bother with path_get()/path_put() in unix_open_file()
fold fs_struct->{lock,seq} into a seqlock
selftests: net: extend SCM_PIDFD test to cover stale pidfds
af_unix: enable handing out pidfds for reaped tasks in SCM_PIDFD
af_unix: stash pidfs dentry when needed
af_unix/scm: fix whitespace errors
af_unix: introduce and use scm_replace_pid() helper
af_unix: introduce unix_skb_to_scm helper
af_unix: rework unix_maybe_add_creds() to allow sleep
selftests/pidfd: decode pidfd file handles withou having to specify an fd
fhandle, pidfs: support open_by_handle_at() purely based on file handle
uapi/fcntl: add FD_PIDFS_ROOT
uapi/fcntl: add FD_INVALID
fcntl/pidfd: redefine PIDFD_SELF_THREAD_GROUP
uapi/fcntl: mark range as reserved
fhandle: reflow get_path_anchor()
pidfs: add pidfs_root_path() helper
fhandle: rename to get_path_anchor()
fhandle: hoist copy_from_user() above get_path_from_fd()
fhandle: raise FILEID_IS_DIR in handle_type
...
This commit is contained in:
@@ -710,11 +710,6 @@ static bool coredump_sock_connect(struct core_name *cn, struct coredump_params *
|
||||
|
||||
retval = kernel_connect(socket, (struct sockaddr *)(&addr), addr_len,
|
||||
O_NONBLOCK | SOCK_COREDUMP);
|
||||
/*
|
||||
* ... Make sure to only put our reference after connect() took
|
||||
* its own reference keeping the pidfs entry alive ...
|
||||
*/
|
||||
pidfs_put_pid(cprm->pid);
|
||||
|
||||
if (retval) {
|
||||
if (retval == -EAGAIN)
|
||||
|
||||
+4
-4
@@ -241,9 +241,9 @@ static void get_fs_root_rcu(struct fs_struct *fs, struct path *root)
|
||||
unsigned seq;
|
||||
|
||||
do {
|
||||
seq = read_seqcount_begin(&fs->seq);
|
||||
seq = read_seqbegin(&fs->seq);
|
||||
*root = fs->root;
|
||||
} while (read_seqcount_retry(&fs->seq, seq));
|
||||
} while (read_seqretry(&fs->seq, seq));
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -385,10 +385,10 @@ static void get_fs_root_and_pwd_rcu(struct fs_struct *fs, struct path *root,
|
||||
unsigned seq;
|
||||
|
||||
do {
|
||||
seq = read_seqcount_begin(&fs->seq);
|
||||
seq = read_seqbegin(&fs->seq);
|
||||
*root = fs->root;
|
||||
*pwd = fs->pwd;
|
||||
} while (read_seqcount_retry(&fs->seq, seq));
|
||||
} while (read_seqretry(&fs->seq, seq));
|
||||
}
|
||||
|
||||
/*
|
||||
|
||||
@@ -1515,7 +1515,7 @@ static void check_unsafe_exec(struct linux_binprm *bprm)
|
||||
* state is protected by cred_guard_mutex we hold.
|
||||
*/
|
||||
n_fs = 1;
|
||||
spin_lock(&p->fs->lock);
|
||||
read_seqlock_excl(&p->fs->seq);
|
||||
rcu_read_lock();
|
||||
for_other_threads(p, t) {
|
||||
if (t->fs == p->fs)
|
||||
@@ -1528,7 +1528,7 @@ static void check_unsafe_exec(struct linux_binprm *bprm)
|
||||
bprm->unsafe |= LSM_UNSAFE_SHARE;
|
||||
else
|
||||
p->fs->in_exec = 1;
|
||||
spin_unlock(&p->fs->lock);
|
||||
read_sequnlock_excl(&p->fs->seq);
|
||||
}
|
||||
|
||||
static void bprm_fill_uid(struct linux_binprm *bprm, struct file *file)
|
||||
|
||||
+30
-32
@@ -88,7 +88,7 @@ static long do_sys_name_to_handle(const struct path *path,
|
||||
if (fh_flags & EXPORT_FH_CONNECTABLE) {
|
||||
handle->handle_type |= FILEID_IS_CONNECTABLE;
|
||||
if (d_is_dir(path->dentry))
|
||||
fh_flags |= FILEID_IS_DIR;
|
||||
handle->handle_type |= FILEID_IS_DIR;
|
||||
}
|
||||
retval = 0;
|
||||
}
|
||||
@@ -168,23 +168,28 @@ SYSCALL_DEFINE5(name_to_handle_at, int, dfd, const char __user *, name,
|
||||
return err;
|
||||
}
|
||||
|
||||
static int get_path_from_fd(int fd, struct path *root)
|
||||
static int get_path_anchor(int fd, struct path *root)
|
||||
{
|
||||
if (fd == AT_FDCWD) {
|
||||
struct fs_struct *fs = current->fs;
|
||||
spin_lock(&fs->lock);
|
||||
*root = fs->pwd;
|
||||
path_get(root);
|
||||
spin_unlock(&fs->lock);
|
||||
} else {
|
||||
if (fd >= 0) {
|
||||
CLASS(fd, f)(fd);
|
||||
if (fd_empty(f))
|
||||
return -EBADF;
|
||||
*root = fd_file(f)->f_path;
|
||||
path_get(root);
|
||||
return 0;
|
||||
}
|
||||
|
||||
return 0;
|
||||
if (fd == AT_FDCWD) {
|
||||
get_fs_pwd(current->fs, root);
|
||||
return 0;
|
||||
}
|
||||
|
||||
if (fd == FD_PIDFS_ROOT) {
|
||||
pidfs_get_root(root);
|
||||
return 0;
|
||||
}
|
||||
|
||||
return -EBADF;
|
||||
}
|
||||
|
||||
static int vfs_dentry_acceptable(void *context, struct dentry *dentry)
|
||||
@@ -323,13 +328,24 @@ static int handle_to_path(int mountdirfd, struct file_handle __user *ufh,
|
||||
{
|
||||
int retval = 0;
|
||||
struct file_handle f_handle;
|
||||
struct file_handle *handle = NULL;
|
||||
struct file_handle *handle __free(kfree) = NULL;
|
||||
struct handle_to_path_ctx ctx = {};
|
||||
const struct export_operations *eops;
|
||||
|
||||
retval = get_path_from_fd(mountdirfd, &ctx.root);
|
||||
if (copy_from_user(&f_handle, ufh, sizeof(struct file_handle)))
|
||||
return -EFAULT;
|
||||
|
||||
if ((f_handle.handle_bytes > MAX_HANDLE_SZ) ||
|
||||
(f_handle.handle_bytes == 0))
|
||||
return -EINVAL;
|
||||
|
||||
if (f_handle.handle_type < 0 ||
|
||||
FILEID_USER_FLAGS(f_handle.handle_type) & ~FILEID_VALID_USER_FLAGS)
|
||||
return -EINVAL;
|
||||
|
||||
retval = get_path_anchor(mountdirfd, &ctx.root);
|
||||
if (retval)
|
||||
goto out_err;
|
||||
return retval;
|
||||
|
||||
eops = ctx.root.mnt->mnt_sb->s_export_op;
|
||||
if (eops && eops->permission)
|
||||
@@ -339,21 +355,6 @@ static int handle_to_path(int mountdirfd, struct file_handle __user *ufh,
|
||||
if (retval)
|
||||
goto out_path;
|
||||
|
||||
if (copy_from_user(&f_handle, ufh, sizeof(struct file_handle))) {
|
||||
retval = -EFAULT;
|
||||
goto out_path;
|
||||
}
|
||||
if ((f_handle.handle_bytes > MAX_HANDLE_SZ) ||
|
||||
(f_handle.handle_bytes == 0)) {
|
||||
retval = -EINVAL;
|
||||
goto out_path;
|
||||
}
|
||||
if (f_handle.handle_type < 0 ||
|
||||
FILEID_USER_FLAGS(f_handle.handle_type) & ~FILEID_VALID_USER_FLAGS) {
|
||||
retval = -EINVAL;
|
||||
goto out_path;
|
||||
}
|
||||
|
||||
handle = kmalloc(struct_size(handle, f_handle, f_handle.handle_bytes),
|
||||
GFP_KERNEL);
|
||||
if (!handle) {
|
||||
@@ -366,7 +367,7 @@ static int handle_to_path(int mountdirfd, struct file_handle __user *ufh,
|
||||
&ufh->f_handle,
|
||||
f_handle.handle_bytes)) {
|
||||
retval = -EFAULT;
|
||||
goto out_handle;
|
||||
goto out_path;
|
||||
}
|
||||
|
||||
/*
|
||||
@@ -384,11 +385,8 @@ static int handle_to_path(int mountdirfd, struct file_handle __user *ufh,
|
||||
handle->handle_type &= ~FILEID_USER_FLAGS_MASK;
|
||||
retval = do_handle_to_path(handle, path, &ctx);
|
||||
|
||||
out_handle:
|
||||
kfree(handle);
|
||||
out_path:
|
||||
path_put(&ctx.root);
|
||||
out_err:
|
||||
return retval;
|
||||
}
|
||||
|
||||
|
||||
+14
-22
@@ -17,12 +17,10 @@ void set_fs_root(struct fs_struct *fs, const struct path *path)
|
||||
struct path old_root;
|
||||
|
||||
path_get(path);
|
||||
spin_lock(&fs->lock);
|
||||
write_seqcount_begin(&fs->seq);
|
||||
write_seqlock(&fs->seq);
|
||||
old_root = fs->root;
|
||||
fs->root = *path;
|
||||
write_seqcount_end(&fs->seq);
|
||||
spin_unlock(&fs->lock);
|
||||
write_sequnlock(&fs->seq);
|
||||
if (old_root.dentry)
|
||||
path_put(&old_root);
|
||||
}
|
||||
@@ -36,12 +34,10 @@ void set_fs_pwd(struct fs_struct *fs, const struct path *path)
|
||||
struct path old_pwd;
|
||||
|
||||
path_get(path);
|
||||
spin_lock(&fs->lock);
|
||||
write_seqcount_begin(&fs->seq);
|
||||
write_seqlock(&fs->seq);
|
||||
old_pwd = fs->pwd;
|
||||
fs->pwd = *path;
|
||||
write_seqcount_end(&fs->seq);
|
||||
spin_unlock(&fs->lock);
|
||||
write_sequnlock(&fs->seq);
|
||||
|
||||
if (old_pwd.dentry)
|
||||
path_put(&old_pwd);
|
||||
@@ -67,16 +63,14 @@ void chroot_fs_refs(const struct path *old_root, const struct path *new_root)
|
||||
fs = p->fs;
|
||||
if (fs) {
|
||||
int hits = 0;
|
||||
spin_lock(&fs->lock);
|
||||
write_seqcount_begin(&fs->seq);
|
||||
write_seqlock(&fs->seq);
|
||||
hits += replace_path(&fs->root, old_root, new_root);
|
||||
hits += replace_path(&fs->pwd, old_root, new_root);
|
||||
write_seqcount_end(&fs->seq);
|
||||
while (hits--) {
|
||||
count++;
|
||||
path_get(new_root);
|
||||
}
|
||||
spin_unlock(&fs->lock);
|
||||
write_sequnlock(&fs->seq);
|
||||
}
|
||||
task_unlock(p);
|
||||
}
|
||||
@@ -99,10 +93,10 @@ void exit_fs(struct task_struct *tsk)
|
||||
if (fs) {
|
||||
int kill;
|
||||
task_lock(tsk);
|
||||
spin_lock(&fs->lock);
|
||||
read_seqlock_excl(&fs->seq);
|
||||
tsk->fs = NULL;
|
||||
kill = !--fs->users;
|
||||
spin_unlock(&fs->lock);
|
||||
read_sequnlock_excl(&fs->seq);
|
||||
task_unlock(tsk);
|
||||
if (kill)
|
||||
free_fs_struct(fs);
|
||||
@@ -116,16 +110,15 @@ struct fs_struct *copy_fs_struct(struct fs_struct *old)
|
||||
if (fs) {
|
||||
fs->users = 1;
|
||||
fs->in_exec = 0;
|
||||
spin_lock_init(&fs->lock);
|
||||
seqcount_spinlock_init(&fs->seq, &fs->lock);
|
||||
seqlock_init(&fs->seq);
|
||||
fs->umask = old->umask;
|
||||
|
||||
spin_lock(&old->lock);
|
||||
read_seqlock_excl(&old->seq);
|
||||
fs->root = old->root;
|
||||
path_get(&fs->root);
|
||||
fs->pwd = old->pwd;
|
||||
path_get(&fs->pwd);
|
||||
spin_unlock(&old->lock);
|
||||
read_sequnlock_excl(&old->seq);
|
||||
}
|
||||
return fs;
|
||||
}
|
||||
@@ -140,10 +133,10 @@ int unshare_fs_struct(void)
|
||||
return -ENOMEM;
|
||||
|
||||
task_lock(current);
|
||||
spin_lock(&fs->lock);
|
||||
read_seqlock_excl(&fs->seq);
|
||||
kill = !--fs->users;
|
||||
current->fs = new_fs;
|
||||
spin_unlock(&fs->lock);
|
||||
read_sequnlock_excl(&fs->seq);
|
||||
task_unlock(current);
|
||||
|
||||
if (kill)
|
||||
@@ -162,7 +155,6 @@ EXPORT_SYMBOL(current_umask);
|
||||
/* to be mentioned only in INIT_TASK */
|
||||
struct fs_struct init_fs = {
|
||||
.users = 1,
|
||||
.lock = __SPIN_LOCK_UNLOCKED(init_fs.lock),
|
||||
.seq = SEQCNT_SPINLOCK_ZERO(init_fs.seq, &init_fs.lock),
|
||||
.seq = __SEQLOCK_UNLOCKED(init_fs.seq),
|
||||
.umask = 0022,
|
||||
};
|
||||
|
||||
@@ -323,12 +323,15 @@ struct mnt_idmap *alloc_mnt_idmap(struct user_namespace *mnt_userns);
|
||||
struct mnt_idmap *mnt_idmap_get(struct mnt_idmap *idmap);
|
||||
void mnt_idmap_put(struct mnt_idmap *idmap);
|
||||
struct stashed_operations {
|
||||
struct dentry *(*stash_dentry)(struct dentry **stashed,
|
||||
struct dentry *dentry);
|
||||
void (*put_data)(void *data);
|
||||
int (*init_inode)(struct inode *inode, void *data);
|
||||
};
|
||||
int path_from_stashed(struct dentry **stashed, struct vfsmount *mnt, void *data,
|
||||
struct path *path);
|
||||
void stashed_dentry_prune(struct dentry *dentry);
|
||||
struct dentry *stash_dentry(struct dentry **stashed, struct dentry *dentry);
|
||||
struct dentry *stashed_dentry_get(struct dentry **stashed);
|
||||
/**
|
||||
* path_mounted - check whether path is mounted
|
||||
@@ -351,3 +354,4 @@ int anon_inode_getattr(struct mnt_idmap *idmap, const struct path *path,
|
||||
unsigned int query_flags);
|
||||
int anon_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry,
|
||||
struct iattr *attr);
|
||||
void pidfs_get_root(struct path *path);
|
||||
|
||||
+22
-12
@@ -2143,6 +2143,8 @@ struct dentry *stashed_dentry_get(struct dentry **stashed)
|
||||
dentry = rcu_dereference(*stashed);
|
||||
if (!dentry)
|
||||
return NULL;
|
||||
if (IS_ERR(dentry))
|
||||
return dentry;
|
||||
if (!lockref_get_not_dead(&dentry->d_lockref))
|
||||
return NULL;
|
||||
return dentry;
|
||||
@@ -2175,7 +2177,6 @@ static struct dentry *prepare_anon_dentry(struct dentry **stashed,
|
||||
|
||||
/* Notice when this is changed. */
|
||||
WARN_ON_ONCE(!S_ISREG(inode->i_mode));
|
||||
WARN_ON_ONCE(!IS_IMMUTABLE(inode));
|
||||
|
||||
dentry = d_alloc_anon(sb);
|
||||
if (!dentry) {
|
||||
@@ -2191,8 +2192,7 @@ static struct dentry *prepare_anon_dentry(struct dentry **stashed,
|
||||
return dentry;
|
||||
}
|
||||
|
||||
static struct dentry *stash_dentry(struct dentry **stashed,
|
||||
struct dentry *dentry)
|
||||
struct dentry *stash_dentry(struct dentry **stashed, struct dentry *dentry)
|
||||
{
|
||||
guard(rcu)();
|
||||
for (;;) {
|
||||
@@ -2233,14 +2233,16 @@ static struct dentry *stash_dentry(struct dentry **stashed,
|
||||
int path_from_stashed(struct dentry **stashed, struct vfsmount *mnt, void *data,
|
||||
struct path *path)
|
||||
{
|
||||
struct dentry *dentry;
|
||||
struct dentry *dentry, *res;
|
||||
const struct stashed_operations *sops = mnt->mnt_sb->s_fs_info;
|
||||
|
||||
/* See if dentry can be reused. */
|
||||
path->dentry = stashed_dentry_get(stashed);
|
||||
if (path->dentry) {
|
||||
res = stashed_dentry_get(stashed);
|
||||
if (IS_ERR(res))
|
||||
return PTR_ERR(res);
|
||||
if (res) {
|
||||
sops->put_data(data);
|
||||
goto out_path;
|
||||
goto make_path;
|
||||
}
|
||||
|
||||
/* Allocate a new dentry. */
|
||||
@@ -2249,14 +2251,22 @@ int path_from_stashed(struct dentry **stashed, struct vfsmount *mnt, void *data,
|
||||
return PTR_ERR(dentry);
|
||||
|
||||
/* Added a new dentry. @data is now owned by the filesystem. */
|
||||
path->dentry = stash_dentry(stashed, dentry);
|
||||
if (path->dentry != dentry)
|
||||
if (sops->stash_dentry)
|
||||
res = sops->stash_dentry(stashed, dentry);
|
||||
else
|
||||
res = stash_dentry(stashed, dentry);
|
||||
if (IS_ERR(res)) {
|
||||
dput(dentry);
|
||||
return PTR_ERR(res);
|
||||
}
|
||||
if (res != dentry)
|
||||
dput(dentry);
|
||||
|
||||
out_path:
|
||||
WARN_ON_ONCE(path->dentry->d_fsdata != stashed);
|
||||
WARN_ON_ONCE(d_inode(path->dentry)->i_private != data);
|
||||
make_path:
|
||||
path->dentry = res;
|
||||
path->mnt = mntget(mnt);
|
||||
VFS_WARN_ON_ONCE(path->dentry->d_fsdata != stashed);
|
||||
VFS_WARN_ON_ONCE(d_inode(path->dentry)->i_private != data);
|
||||
return 0;
|
||||
}
|
||||
|
||||
|
||||
+4
-4
@@ -1012,10 +1012,10 @@ static int set_root(struct nameidata *nd)
|
||||
unsigned seq;
|
||||
|
||||
do {
|
||||
seq = read_seqcount_begin(&fs->seq);
|
||||
seq = read_seqbegin(&fs->seq);
|
||||
nd->root = fs->root;
|
||||
nd->root_seq = __read_seqcount_begin(&nd->root.dentry->d_seq);
|
||||
} while (read_seqcount_retry(&fs->seq, seq));
|
||||
} while (read_seqretry(&fs->seq, seq));
|
||||
} else {
|
||||
get_fs_root(fs, &nd->root);
|
||||
nd->state |= ND_ROOT_GRABBED;
|
||||
@@ -2571,11 +2571,11 @@ static const char *path_init(struct nameidata *nd, unsigned flags)
|
||||
unsigned seq;
|
||||
|
||||
do {
|
||||
seq = read_seqcount_begin(&fs->seq);
|
||||
seq = read_seqbegin(&fs->seq);
|
||||
nd->path = fs->pwd;
|
||||
nd->inode = nd->path.dentry->d_inode;
|
||||
nd->seq = __read_seqcount_begin(&nd->path.dentry->d_seq);
|
||||
} while (read_seqcount_retry(&fs->seq, seq));
|
||||
} while (read_seqretry(&fs->seq, seq));
|
||||
} else {
|
||||
get_fs_pwd(current->fs, &nd->path);
|
||||
nd->inode = nd->path.dentry->d_inode;
|
||||
|
||||
+255
-185
File diff suppressed because it is too large
Load Diff
@@ -8,8 +8,7 @@
|
||||
|
||||
struct fs_struct {
|
||||
int users;
|
||||
spinlock_t lock;
|
||||
seqcount_spinlock_t seq;
|
||||
seqlock_t seq;
|
||||
int umask;
|
||||
int in_exec;
|
||||
struct path root, pwd;
|
||||
@@ -26,18 +25,18 @@ extern int unshare_fs_struct(void);
|
||||
|
||||
static inline void get_fs_root(struct fs_struct *fs, struct path *root)
|
||||
{
|
||||
spin_lock(&fs->lock);
|
||||
read_seqlock_excl(&fs->seq);
|
||||
*root = fs->root;
|
||||
path_get(root);
|
||||
spin_unlock(&fs->lock);
|
||||
read_sequnlock_excl(&fs->seq);
|
||||
}
|
||||
|
||||
static inline void get_fs_pwd(struct fs_struct *fs, struct path *pwd)
|
||||
{
|
||||
spin_lock(&fs->lock);
|
||||
read_seqlock_excl(&fs->seq);
|
||||
*pwd = fs->pwd;
|
||||
path_get(pwd);
|
||||
spin_unlock(&fs->lock);
|
||||
read_sequnlock_excl(&fs->seq);
|
||||
}
|
||||
|
||||
extern bool current_chrooted(void);
|
||||
|
||||
+9
-5
@@ -47,19 +47,23 @@
|
||||
|
||||
#define RESERVED_PIDS 300
|
||||
|
||||
struct pidfs_attr;
|
||||
|
||||
struct upid {
|
||||
int nr;
|
||||
struct pid_namespace *ns;
|
||||
};
|
||||
|
||||
struct pid
|
||||
{
|
||||
struct pid {
|
||||
refcount_t count;
|
||||
unsigned int level;
|
||||
spinlock_t lock;
|
||||
struct dentry *stashed;
|
||||
u64 ino;
|
||||
struct rb_node pidfs_node;
|
||||
struct {
|
||||
u64 ino;
|
||||
struct rb_node pidfs_node;
|
||||
struct dentry *stashed;
|
||||
struct pidfs_attr *attr;
|
||||
};
|
||||
/* lists of tasks that use this pid */
|
||||
struct hlist_head tasks[PIDTYPE_MAX];
|
||||
struct hlist_head inodes;
|
||||
|
||||
@@ -14,7 +14,6 @@ void pidfs_coredump(const struct coredump_params *cprm);
|
||||
#endif
|
||||
extern const struct dentry_operations pidfs_dentry_operations;
|
||||
int pidfs_register_pid(struct pid *pid);
|
||||
void pidfs_get_pid(struct pid *pid);
|
||||
void pidfs_put_pid(struct pid *pid);
|
||||
void pidfs_free_pid(struct pid *pid);
|
||||
|
||||
#endif /* _LINUX_PID_FS_H */
|
||||
|
||||
+2
-2
@@ -69,7 +69,7 @@ static __inline__ void unix_get_peersec_dgram(struct socket *sock, struct scm_co
|
||||
static __inline__ void scm_set_cred(struct scm_cookie *scm,
|
||||
struct pid *pid, kuid_t uid, kgid_t gid)
|
||||
{
|
||||
scm->pid = get_pid(pid);
|
||||
scm->pid = get_pid(pid);
|
||||
scm->creds.pid = pid_vnr(pid);
|
||||
scm->creds.uid = uid;
|
||||
scm->creds.gid = gid;
|
||||
@@ -78,7 +78,7 @@ static __inline__ void scm_set_cred(struct scm_cookie *scm,
|
||||
static __inline__ void scm_destroy_cred(struct scm_cookie *scm)
|
||||
{
|
||||
put_pid(scm->pid);
|
||||
scm->pid = NULL;
|
||||
scm->pid = NULL;
|
||||
}
|
||||
|
||||
static __inline__ void scm_destroy(struct scm_cookie *scm)
|
||||
|
||||
@@ -90,10 +90,28 @@
|
||||
#define DN_ATTRIB 0x00000020 /* File changed attibutes */
|
||||
#define DN_MULTISHOT 0x80000000 /* Don't remove notifier */
|
||||
|
||||
/* Reserved kernel ranges [-100], [-10000, -40000]. */
|
||||
#define AT_FDCWD -100 /* Special value for dirfd used to
|
||||
indicate openat should use the
|
||||
current working directory. */
|
||||
|
||||
/*
|
||||
* The concept of process and threads in userland and the kernel is a confusing
|
||||
* one - within the kernel every thread is a 'task' with its own individual PID,
|
||||
* however from userland's point of view threads are grouped by a single PID,
|
||||
* which is that of the 'thread group leader', typically the first thread
|
||||
* spawned.
|
||||
*
|
||||
* To cut the Gideon knot, for internal kernel usage, we refer to
|
||||
* PIDFD_SELF_THREAD to refer to the current thread (or task from a kernel
|
||||
* perspective), and PIDFD_SELF_THREAD_GROUP to refer to the current thread
|
||||
* group leader...
|
||||
*/
|
||||
#define PIDFD_SELF_THREAD -10000 /* Current thread. */
|
||||
#define PIDFD_SELF_THREAD_GROUP -10001 /* Current thread group leader. */
|
||||
|
||||
#define FD_PIDFS_ROOT -10002 /* Root of the pidfs filesystem */
|
||||
#define FD_INVALID -10009 /* Invalid file descriptor: -10000 - EBADF = -10009 */
|
||||
|
||||
/* Generic flags for the *at(2) family of syscalls. */
|
||||
|
||||
|
||||
@@ -42,21 +42,6 @@
|
||||
#define PIDFD_COREDUMP_USER (1U << 2) /* coredump was done as the user. */
|
||||
#define PIDFD_COREDUMP_ROOT (1U << 3) /* coredump was done as root. */
|
||||
|
||||
/*
|
||||
* The concept of process and threads in userland and the kernel is a confusing
|
||||
* one - within the kernel every thread is a 'task' with its own individual PID,
|
||||
* however from userland's point of view threads are grouped by a single PID,
|
||||
* which is that of the 'thread group leader', typically the first thread
|
||||
* spawned.
|
||||
*
|
||||
* To cut the Gideon knot, for internal kernel usage, we refer to
|
||||
* PIDFD_SELF_THREAD to refer to the current thread (or task from a kernel
|
||||
* perspective), and PIDFD_SELF_THREAD_GROUP to refer to the current thread
|
||||
* group leader...
|
||||
*/
|
||||
#define PIDFD_SELF_THREAD -10000 /* Current thread. */
|
||||
#define PIDFD_SELF_THREAD_GROUP -20000 /* Current thread group leader. */
|
||||
|
||||
/*
|
||||
* ...and for userland we make life simpler - PIDFD_SELF refers to the current
|
||||
* thread, PIDFD_SELF_PROCESS refers to the process thread group leader.
|
||||
|
||||
+5
-5
@@ -1542,14 +1542,14 @@ static int copy_fs(unsigned long clone_flags, struct task_struct *tsk)
|
||||
struct fs_struct *fs = current->fs;
|
||||
if (clone_flags & CLONE_FS) {
|
||||
/* tsk->fs is already what we want */
|
||||
spin_lock(&fs->lock);
|
||||
read_seqlock_excl(&fs->seq);
|
||||
/* "users" and "in_exec" locked for check_unsafe_exec() */
|
||||
if (fs->in_exec) {
|
||||
spin_unlock(&fs->lock);
|
||||
read_sequnlock_excl(&fs->seq);
|
||||
return -EAGAIN;
|
||||
}
|
||||
fs->users++;
|
||||
spin_unlock(&fs->lock);
|
||||
read_sequnlock_excl(&fs->seq);
|
||||
return 0;
|
||||
}
|
||||
tsk->fs = copy_fs_struct(fs);
|
||||
@@ -3149,13 +3149,13 @@ int ksys_unshare(unsigned long unshare_flags)
|
||||
|
||||
if (new_fs) {
|
||||
fs = current->fs;
|
||||
spin_lock(&fs->lock);
|
||||
read_seqlock_excl(&fs->seq);
|
||||
current->fs = new_fs;
|
||||
if (--fs->users)
|
||||
new_fs = NULL;
|
||||
else
|
||||
new_fs = fs;
|
||||
spin_unlock(&fs->lock);
|
||||
read_sequnlock_excl(&fs->seq);
|
||||
}
|
||||
|
||||
if (new_fd)
|
||||
|
||||
+1
-1
@@ -100,7 +100,7 @@ void put_pid(struct pid *pid)
|
||||
|
||||
ns = pid->numbers[pid->level].ns;
|
||||
if (refcount_dec_and_test(&pid->count)) {
|
||||
WARN_ON_ONCE(pid->stashed);
|
||||
pidfs_free_pid(pid);
|
||||
kmem_cache_free(ns->pid_cachep, pid);
|
||||
put_pid_ns(ns);
|
||||
}
|
||||
|
||||
+28
-4
@@ -23,6 +23,8 @@
|
||||
#include <linux/security.h>
|
||||
#include <linux/pid_namespace.h>
|
||||
#include <linux/pid.h>
|
||||
#include <uapi/linux/pidfd.h>
|
||||
#include <linux/pidfs.h>
|
||||
#include <linux/nsproxy.h>
|
||||
#include <linux/slab.h>
|
||||
#include <linux/errqueue.h>
|
||||
@@ -145,6 +147,22 @@ void __scm_destroy(struct scm_cookie *scm)
|
||||
}
|
||||
EXPORT_SYMBOL(__scm_destroy);
|
||||
|
||||
static inline int scm_replace_pid(struct scm_cookie *scm, struct pid *pid)
|
||||
{
|
||||
int err;
|
||||
|
||||
/* drop all previous references */
|
||||
scm_destroy_cred(scm);
|
||||
|
||||
err = pidfs_register_pid(pid);
|
||||
if (unlikely(err))
|
||||
return err;
|
||||
|
||||
scm->pid = pid;
|
||||
scm->creds.pid = pid_vnr(pid);
|
||||
return 0;
|
||||
}
|
||||
|
||||
int __scm_send(struct socket *sock, struct msghdr *msg, struct scm_cookie *p)
|
||||
{
|
||||
const struct proto_ops *ops = READ_ONCE(sock->ops);
|
||||
@@ -189,15 +207,21 @@ int __scm_send(struct socket *sock, struct msghdr *msg, struct scm_cookie *p)
|
||||
if (err)
|
||||
goto error;
|
||||
|
||||
p->creds.pid = creds.pid;
|
||||
if (!p->pid || pid_vnr(p->pid) != creds.pid) {
|
||||
struct pid *pid;
|
||||
err = -ESRCH;
|
||||
pid = find_get_pid(creds.pid);
|
||||
if (!pid)
|
||||
goto error;
|
||||
put_pid(p->pid);
|
||||
p->pid = pid;
|
||||
|
||||
/* pass a struct pid reference from
|
||||
* find_get_pid() to scm_replace_pid().
|
||||
*/
|
||||
err = scm_replace_pid(p, pid);
|
||||
if (err) {
|
||||
put_pid(pid);
|
||||
goto error;
|
||||
}
|
||||
}
|
||||
|
||||
err = -EINVAL;
|
||||
@@ -459,7 +483,7 @@ static void scm_pidfd_recv(struct msghdr *msg, struct scm_cookie *scm)
|
||||
if (!scm->pid)
|
||||
return;
|
||||
|
||||
pidfd = pidfd_prepare(scm->pid, 0, &pidfd_file);
|
||||
pidfd = pidfd_prepare(scm->pid, PIDFD_STALE, &pidfd_file);
|
||||
|
||||
if (put_cmsg(msg, SOL_SOCKET, SCM_PIDFD, sizeof(int), &pidfd)) {
|
||||
if (pidfd_file) {
|
||||
|
||||
+48
-30
@@ -646,9 +646,6 @@ static void unix_sock_destructor(struct sock *sk)
|
||||
return;
|
||||
}
|
||||
|
||||
if (sk->sk_peer_pid)
|
||||
pidfs_put_pid(sk->sk_peer_pid);
|
||||
|
||||
if (u->addr)
|
||||
unix_release_addr(u->addr);
|
||||
|
||||
@@ -780,7 +777,6 @@ static void drop_peercred(struct unix_peercred *peercred)
|
||||
swap(peercred->peer_pid, pid);
|
||||
swap(peercred->peer_cred, cred);
|
||||
|
||||
pidfs_put_pid(pid);
|
||||
put_pid(pid);
|
||||
put_cred(cred);
|
||||
}
|
||||
@@ -813,7 +809,6 @@ static void copy_peercred(struct sock *sk, struct sock *peersk)
|
||||
|
||||
spin_lock(&sk->sk_peer_lock);
|
||||
sk->sk_peer_pid = get_pid(peersk->sk_peer_pid);
|
||||
pidfs_get_pid(sk->sk_peer_pid);
|
||||
sk->sk_peer_cred = get_cred(peersk->sk_peer_cred);
|
||||
spin_unlock(&sk->sk_peer_lock);
|
||||
}
|
||||
@@ -1945,7 +1940,7 @@ static void unix_destruct_scm(struct sk_buff *skb)
|
||||
struct scm_cookie scm;
|
||||
|
||||
memset(&scm, 0, sizeof(scm));
|
||||
scm.pid = UNIXCB(skb).pid;
|
||||
scm.pid = UNIXCB(skb).pid;
|
||||
if (UNIXCB(skb).fp)
|
||||
unix_detach_fds(&scm, skb);
|
||||
|
||||
@@ -1971,22 +1966,46 @@ static int unix_scm_to_skb(struct scm_cookie *scm, struct sk_buff *skb, bool sen
|
||||
return err;
|
||||
}
|
||||
|
||||
/*
|
||||
static void unix_skb_to_scm(struct sk_buff *skb, struct scm_cookie *scm)
|
||||
{
|
||||
scm_set_cred(scm, UNIXCB(skb).pid, UNIXCB(skb).uid, UNIXCB(skb).gid);
|
||||
unix_set_secdata(scm, skb);
|
||||
}
|
||||
|
||||
/**
|
||||
* unix_maybe_add_creds() - Adds current task uid/gid and struct pid to skb if needed.
|
||||
* @skb: skb to attach creds to.
|
||||
* @sk: Sender sock.
|
||||
* @other: Receiver sock.
|
||||
*
|
||||
* Some apps rely on write() giving SCM_CREDENTIALS
|
||||
* We include credentials if source or destination socket
|
||||
* asserted SOCK_PASSCRED.
|
||||
*
|
||||
* Context: May sleep.
|
||||
* Return: On success zero, on error a negative error code is returned.
|
||||
*/
|
||||
static void unix_maybe_add_creds(struct sk_buff *skb, const struct sock *sk,
|
||||
const struct sock *other)
|
||||
static int unix_maybe_add_creds(struct sk_buff *skb, const struct sock *sk,
|
||||
const struct sock *other)
|
||||
{
|
||||
if (UNIXCB(skb).pid)
|
||||
return;
|
||||
return 0;
|
||||
|
||||
if (unix_may_passcred(sk) || unix_may_passcred(other) ||
|
||||
!other->sk_socket) {
|
||||
UNIXCB(skb).pid = get_pid(task_tgid(current));
|
||||
struct pid *pid;
|
||||
int err;
|
||||
|
||||
pid = task_tgid(current);
|
||||
err = pidfs_register_pid(pid);
|
||||
if (unlikely(err))
|
||||
return err;
|
||||
|
||||
UNIXCB(skb).pid = get_pid(pid);
|
||||
current_uid_gid(&UNIXCB(skb).uid, &UNIXCB(skb).gid);
|
||||
}
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
static bool unix_skb_scm_eq(struct sk_buff *skb,
|
||||
@@ -2121,6 +2140,10 @@ lookup:
|
||||
goto out_sock_put;
|
||||
}
|
||||
|
||||
err = unix_maybe_add_creds(skb, sk, other);
|
||||
if (err)
|
||||
goto out_sock_put;
|
||||
|
||||
restart:
|
||||
sk_locked = 0;
|
||||
unix_state_lock(other);
|
||||
@@ -2229,7 +2252,6 @@ restart_locked:
|
||||
if (sock_flag(other, SOCK_RCVTSTAMP))
|
||||
__net_timestamp(skb);
|
||||
|
||||
unix_maybe_add_creds(skb, sk, other);
|
||||
scm_stat_add(other, skb);
|
||||
skb_queue_tail(&other->sk_receive_queue, skb);
|
||||
unix_state_unlock(other);
|
||||
@@ -2273,6 +2295,10 @@ static int queue_oob(struct sock *sk, struct msghdr *msg, struct sock *other,
|
||||
if (err < 0)
|
||||
goto out;
|
||||
|
||||
err = unix_maybe_add_creds(skb, sk, other);
|
||||
if (err)
|
||||
goto out;
|
||||
|
||||
skb_put(skb, 1);
|
||||
err = skb_copy_datagram_from_iter(skb, 0, &msg->msg_iter, 1);
|
||||
|
||||
@@ -2292,7 +2318,6 @@ static int queue_oob(struct sock *sk, struct msghdr *msg, struct sock *other,
|
||||
goto out_unlock;
|
||||
}
|
||||
|
||||
unix_maybe_add_creds(skb, sk, other);
|
||||
scm_stat_add(other, skb);
|
||||
|
||||
spin_lock(&other->sk_receive_queue.lock);
|
||||
@@ -2386,6 +2411,10 @@ static int unix_stream_sendmsg(struct socket *sock, struct msghdr *msg,
|
||||
|
||||
fds_sent = true;
|
||||
|
||||
err = unix_maybe_add_creds(skb, sk, other);
|
||||
if (err)
|
||||
goto out_free;
|
||||
|
||||
if (unlikely(msg->msg_flags & MSG_SPLICE_PAGES)) {
|
||||
skb->ip_summed = CHECKSUM_UNNECESSARY;
|
||||
err = skb_splice_from_iter(skb, &msg->msg_iter, size,
|
||||
@@ -2416,7 +2445,6 @@ static int unix_stream_sendmsg(struct socket *sock, struct msghdr *msg,
|
||||
goto out_free;
|
||||
}
|
||||
|
||||
unix_maybe_add_creds(skb, sk, other);
|
||||
scm_stat_add(other, skb);
|
||||
skb_queue_tail(&other->sk_receive_queue, skb);
|
||||
unix_state_unlock(other);
|
||||
@@ -2564,8 +2592,7 @@ int __unix_dgram_recvmsg(struct sock *sk, struct msghdr *msg, size_t size,
|
||||
|
||||
memset(&scm, 0, sizeof(scm));
|
||||
|
||||
scm_set_cred(&scm, UNIXCB(skb).pid, UNIXCB(skb).uid, UNIXCB(skb).gid);
|
||||
unix_set_secdata(&scm, skb);
|
||||
unix_skb_to_scm(skb, &scm);
|
||||
|
||||
if (!(flags & MSG_PEEK)) {
|
||||
if (UNIXCB(skb).fp)
|
||||
@@ -2954,8 +2981,7 @@ unlock:
|
||||
break;
|
||||
} else if (unix_may_passcred(sk)) {
|
||||
/* Copy credentials */
|
||||
scm_set_cred(&scm, UNIXCB(skb).pid, UNIXCB(skb).uid, UNIXCB(skb).gid);
|
||||
unix_set_secdata(&scm, skb);
|
||||
unix_skb_to_scm(skb, &scm);
|
||||
check_creds = true;
|
||||
}
|
||||
|
||||
@@ -3191,7 +3217,6 @@ EXPORT_SYMBOL_GPL(unix_outq_len);
|
||||
|
||||
static int unix_open_file(struct sock *sk)
|
||||
{
|
||||
struct path path;
|
||||
struct file *f;
|
||||
int fd;
|
||||
|
||||
@@ -3201,27 +3226,20 @@ static int unix_open_file(struct sock *sk)
|
||||
if (!smp_load_acquire(&unix_sk(sk)->addr))
|
||||
return -ENOENT;
|
||||
|
||||
path = unix_sk(sk)->path;
|
||||
if (!path.dentry)
|
||||
if (!unix_sk(sk)->path.dentry)
|
||||
return -ENOENT;
|
||||
|
||||
path_get(&path);
|
||||
|
||||
fd = get_unused_fd_flags(O_CLOEXEC);
|
||||
if (fd < 0)
|
||||
goto out;
|
||||
return fd;
|
||||
|
||||
f = dentry_open(&path, O_PATH, current_cred());
|
||||
f = dentry_open(&unix_sk(sk)->path, O_PATH, current_cred());
|
||||
if (IS_ERR(f)) {
|
||||
put_unused_fd(fd);
|
||||
fd = PTR_ERR(f);
|
||||
goto out;
|
||||
return PTR_ERR(f);
|
||||
}
|
||||
|
||||
fd_install(fd, f);
|
||||
out:
|
||||
path_put(&path);
|
||||
|
||||
return fd;
|
||||
}
|
||||
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
#include <sys/types.h>
|
||||
#include <sys/wait.h>
|
||||
|
||||
#include "../../pidfd/pidfd.h"
|
||||
#include "../../kselftest_harness.h"
|
||||
|
||||
#define clean_errno() (errno == 0 ? "None" : strerror(errno))
|
||||
@@ -26,6 +27,8 @@
|
||||
#define SCM_PIDFD 0x04
|
||||
#endif
|
||||
|
||||
#define CHILD_EXIT_CODE_OK 123
|
||||
|
||||
static void child_die()
|
||||
{
|
||||
exit(1);
|
||||
@@ -126,16 +129,65 @@ out:
|
||||
return result;
|
||||
}
|
||||
|
||||
struct cmsg_data {
|
||||
struct ucred *ucred;
|
||||
int *pidfd;
|
||||
};
|
||||
|
||||
static int parse_cmsg(struct msghdr *msg, struct cmsg_data *res)
|
||||
{
|
||||
struct cmsghdr *cmsg;
|
||||
int data = 0;
|
||||
|
||||
if (msg->msg_flags & (MSG_TRUNC | MSG_CTRUNC)) {
|
||||
log_err("recvmsg: truncated");
|
||||
return 1;
|
||||
}
|
||||
|
||||
for (cmsg = CMSG_FIRSTHDR(msg); cmsg != NULL;
|
||||
cmsg = CMSG_NXTHDR(msg, cmsg)) {
|
||||
if (cmsg->cmsg_level == SOL_SOCKET &&
|
||||
cmsg->cmsg_type == SCM_PIDFD) {
|
||||
if (cmsg->cmsg_len < sizeof(*res->pidfd)) {
|
||||
log_err("CMSG parse: SCM_PIDFD wrong len");
|
||||
return 1;
|
||||
}
|
||||
|
||||
res->pidfd = (void *)CMSG_DATA(cmsg);
|
||||
}
|
||||
|
||||
if (cmsg->cmsg_level == SOL_SOCKET &&
|
||||
cmsg->cmsg_type == SCM_CREDENTIALS) {
|
||||
if (cmsg->cmsg_len < sizeof(*res->ucred)) {
|
||||
log_err("CMSG parse: SCM_CREDENTIALS wrong len");
|
||||
return 1;
|
||||
}
|
||||
|
||||
res->ucred = (void *)CMSG_DATA(cmsg);
|
||||
}
|
||||
}
|
||||
|
||||
if (!res->pidfd) {
|
||||
log_err("CMSG parse: SCM_PIDFD not found");
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (!res->ucred) {
|
||||
log_err("CMSG parse: SCM_CREDENTIALS not found");
|
||||
return 1;
|
||||
}
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int cmsg_check(int fd)
|
||||
{
|
||||
struct msghdr msg = { 0 };
|
||||
struct cmsghdr *cmsg;
|
||||
struct cmsg_data res;
|
||||
struct iovec iov;
|
||||
struct ucred *ucred = NULL;
|
||||
int data = 0;
|
||||
char control[CMSG_SPACE(sizeof(struct ucred)) +
|
||||
CMSG_SPACE(sizeof(int))] = { 0 };
|
||||
int *pidfd = NULL;
|
||||
pid_t parent_pid;
|
||||
int err;
|
||||
|
||||
@@ -158,53 +210,99 @@ static int cmsg_check(int fd)
|
||||
return 1;
|
||||
}
|
||||
|
||||
for (cmsg = CMSG_FIRSTHDR(&msg); cmsg != NULL;
|
||||
cmsg = CMSG_NXTHDR(&msg, cmsg)) {
|
||||
if (cmsg->cmsg_level == SOL_SOCKET &&
|
||||
cmsg->cmsg_type == SCM_PIDFD) {
|
||||
if (cmsg->cmsg_len < sizeof(*pidfd)) {
|
||||
log_err("CMSG parse: SCM_PIDFD wrong len");
|
||||
return 1;
|
||||
}
|
||||
|
||||
pidfd = (void *)CMSG_DATA(cmsg);
|
||||
}
|
||||
|
||||
if (cmsg->cmsg_level == SOL_SOCKET &&
|
||||
cmsg->cmsg_type == SCM_CREDENTIALS) {
|
||||
if (cmsg->cmsg_len < sizeof(*ucred)) {
|
||||
log_err("CMSG parse: SCM_CREDENTIALS wrong len");
|
||||
return 1;
|
||||
}
|
||||
|
||||
ucred = (void *)CMSG_DATA(cmsg);
|
||||
}
|
||||
}
|
||||
|
||||
/* send(pfd, "x", sizeof(char), 0) */
|
||||
if (data != 'x') {
|
||||
log_err("recvmsg: data corruption");
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (!pidfd) {
|
||||
log_err("CMSG parse: SCM_PIDFD not found");
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (!ucred) {
|
||||
log_err("CMSG parse: SCM_CREDENTIALS not found");
|
||||
if (parse_cmsg(&msg, &res)) {
|
||||
log_err("CMSG parse: parse_cmsg() failed");
|
||||
return 1;
|
||||
}
|
||||
|
||||
/* pidfd from SCM_PIDFD should point to the parent process PID */
|
||||
parent_pid =
|
||||
get_pid_from_fdinfo_file(*pidfd, "Pid:", sizeof("Pid:") - 1);
|
||||
get_pid_from_fdinfo_file(*res.pidfd, "Pid:", sizeof("Pid:") - 1);
|
||||
if (parent_pid != getppid()) {
|
||||
log_err("wrong SCM_PIDFD %d != %d", parent_pid, getppid());
|
||||
close(*res.pidfd);
|
||||
return 1;
|
||||
}
|
||||
|
||||
close(*res.pidfd);
|
||||
return 0;
|
||||
}
|
||||
|
||||
static int cmsg_check_dead(int fd, int expected_pid)
|
||||
{
|
||||
int err;
|
||||
struct msghdr msg = { 0 };
|
||||
struct cmsg_data res;
|
||||
struct iovec iov;
|
||||
int data = 0;
|
||||
char control[CMSG_SPACE(sizeof(struct ucred)) +
|
||||
CMSG_SPACE(sizeof(int))] = { 0 };
|
||||
pid_t client_pid;
|
||||
struct pidfd_info info = {
|
||||
.mask = PIDFD_INFO_EXIT,
|
||||
};
|
||||
|
||||
iov.iov_base = &data;
|
||||
iov.iov_len = sizeof(data);
|
||||
|
||||
msg.msg_iov = &iov;
|
||||
msg.msg_iovlen = 1;
|
||||
msg.msg_control = control;
|
||||
msg.msg_controllen = sizeof(control);
|
||||
|
||||
err = recvmsg(fd, &msg, 0);
|
||||
if (err < 0) {
|
||||
log_err("recvmsg");
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (msg.msg_flags & (MSG_TRUNC | MSG_CTRUNC)) {
|
||||
log_err("recvmsg: truncated");
|
||||
return 1;
|
||||
}
|
||||
|
||||
/* send(cfd, "y", sizeof(char), 0) */
|
||||
if (data != 'y') {
|
||||
log_err("recvmsg: data corruption");
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (parse_cmsg(&msg, &res)) {
|
||||
log_err("CMSG parse: parse_cmsg() failed");
|
||||
return 1;
|
||||
}
|
||||
|
||||
/*
|
||||
* pidfd from SCM_PIDFD should point to the client_pid.
|
||||
* Let's read exit information and check if it's what
|
||||
* we expect to see.
|
||||
*/
|
||||
if (ioctl(*res.pidfd, PIDFD_GET_INFO, &info)) {
|
||||
log_err("%s: ioctl(PIDFD_GET_INFO) failed", __func__);
|
||||
close(*res.pidfd);
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (!(info.mask & PIDFD_INFO_EXIT)) {
|
||||
log_err("%s: No exit information from ioctl(PIDFD_GET_INFO)", __func__);
|
||||
close(*res.pidfd);
|
||||
return 1;
|
||||
}
|
||||
|
||||
err = WIFEXITED(info.exit_code) ? WEXITSTATUS(info.exit_code) : 1;
|
||||
if (err != CHILD_EXIT_CODE_OK) {
|
||||
log_err("%s: wrong exit_code %d != %d", __func__, err, CHILD_EXIT_CODE_OK);
|
||||
close(*res.pidfd);
|
||||
return 1;
|
||||
}
|
||||
|
||||
close(*res.pidfd);
|
||||
return 0;
|
||||
}
|
||||
|
||||
@@ -291,6 +389,24 @@ static void fill_sockaddr(struct sock_addr *addr, bool abstract)
|
||||
memcpy(sun_path_buf, addr->sock_name, strlen(addr->sock_name));
|
||||
}
|
||||
|
||||
static int sk_enable_cred_pass(int sk)
|
||||
{
|
||||
int on = 0;
|
||||
|
||||
on = 1;
|
||||
if (setsockopt(sk, SOL_SOCKET, SO_PASSCRED, &on, sizeof(on))) {
|
||||
log_err("Failed to set SO_PASSCRED");
|
||||
return 1;
|
||||
}
|
||||
|
||||
if (setsockopt(sk, SOL_SOCKET, SO_PASSPIDFD, &on, sizeof(on))) {
|
||||
log_err("Failed to set SO_PASSPIDFD");
|
||||
return 1;
|
||||
}
|
||||
|
||||
return 0;
|
||||
}
|
||||
|
||||
static void client(FIXTURE_DATA(scm_pidfd) *self,
|
||||
const FIXTURE_VARIANT(scm_pidfd) *variant)
|
||||
{
|
||||
@@ -299,7 +415,6 @@ static void client(FIXTURE_DATA(scm_pidfd) *self,
|
||||
struct ucred peer_cred;
|
||||
int peer_pidfd;
|
||||
pid_t peer_pid;
|
||||
int on = 0;
|
||||
|
||||
cfd = socket(AF_UNIX, variant->type, 0);
|
||||
if (cfd < 0) {
|
||||
@@ -322,14 +437,8 @@ static void client(FIXTURE_DATA(scm_pidfd) *self,
|
||||
child_die();
|
||||
}
|
||||
|
||||
on = 1;
|
||||
if (setsockopt(cfd, SOL_SOCKET, SO_PASSCRED, &on, sizeof(on))) {
|
||||
log_err("Failed to set SO_PASSCRED");
|
||||
child_die();
|
||||
}
|
||||
|
||||
if (setsockopt(cfd, SOL_SOCKET, SO_PASSPIDFD, &on, sizeof(on))) {
|
||||
log_err("Failed to set SO_PASSPIDFD");
|
||||
if (sk_enable_cred_pass(cfd)) {
|
||||
log_err("sk_enable_cred_pass() failed");
|
||||
child_die();
|
||||
}
|
||||
|
||||
@@ -340,6 +449,12 @@ static void client(FIXTURE_DATA(scm_pidfd) *self,
|
||||
child_die();
|
||||
}
|
||||
|
||||
/* send something to the parent so it can receive SCM_PIDFD too and validate it */
|
||||
if (send(cfd, "y", sizeof(char), 0) == -1) {
|
||||
log_err("Failed to send(cfd, \"y\", sizeof(char), 0)");
|
||||
child_die();
|
||||
}
|
||||
|
||||
/* skip further for SOCK_DGRAM as it's not applicable */
|
||||
if (variant->type == SOCK_DGRAM)
|
||||
return;
|
||||
@@ -398,7 +513,13 @@ TEST_F(scm_pidfd, test)
|
||||
close(self->server);
|
||||
close(self->startup_pipe[0]);
|
||||
client(self, variant);
|
||||
exit(0);
|
||||
|
||||
/*
|
||||
* It's a bit unusual, but in case of success we return non-zero
|
||||
* exit code (CHILD_EXIT_CODE_OK) and then we expect to read it
|
||||
* from ioctl(PIDFD_GET_INFO) in cmsg_check_dead().
|
||||
*/
|
||||
exit(CHILD_EXIT_CODE_OK);
|
||||
}
|
||||
close(self->startup_pipe[1]);
|
||||
|
||||
@@ -421,9 +542,17 @@ TEST_F(scm_pidfd, test)
|
||||
ASSERT_NE(-1, err);
|
||||
}
|
||||
|
||||
close(pfd);
|
||||
waitpid(self->client_pid, &child_status, 0);
|
||||
ASSERT_EQ(0, WIFEXITED(child_status) ? WEXITSTATUS(child_status) : 1);
|
||||
/* see comment before exit(CHILD_EXIT_CODE_OK) */
|
||||
ASSERT_EQ(CHILD_EXIT_CODE_OK, WIFEXITED(child_status) ? WEXITSTATUS(child_status) : 1);
|
||||
|
||||
err = sk_enable_cred_pass(pfd);
|
||||
ASSERT_EQ(0, err);
|
||||
|
||||
err = cmsg_check_dead(pfd, self->client_pid);
|
||||
ASSERT_EQ(0, err);
|
||||
|
||||
close(pfd);
|
||||
}
|
||||
|
||||
TEST_HARNESS_MAIN
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user