2019-06-04 10:11:33 +02:00
// SPDX-License-Identifier: GPL-2.0-only
2005-04-16 15:20:36 -07:00
/*
* linux / arch / arm / kernel / signal . c
*
2009-10-25 15:39:37 +00:00
* Copyright ( C ) 1995 - 2009 Russell King
2005-04-16 15:20:36 -07:00
*/
# include <linux/errno.h>
2013-07-24 00:29:18 +01:00
# include <linux/random.h>
2005-04-16 15:20:36 -07:00
# include <linux/signal.h>
# include <linux/personality.h>
2008-09-06 11:35:55 +01:00
# include <linux/uaccess.h>
2022-02-09 12:20:45 -06:00
# include <linux/resume_user_mode.h>
2014-03-07 11:23:04 -05:00
# include <linux/uprobes.h>
2017-09-07 08:30:46 -07:00
# include <linux/syscalls.h>
2005-04-16 15:20:36 -07:00
2006-11-09 14:20:47 +00:00
# include <asm/elf.h>
2005-04-16 15:20:36 -07:00
# include <asm/cacheflush.h>
2013-07-24 00:29:18 +01:00
# include <asm/traps.h>
2005-04-16 15:20:36 -07:00
# include <asm/unistd.h>
2010-04-11 15:58:27 +01:00
# include <asm/vfp.h>
2005-04-16 15:20:36 -07:00
2017-08-09 23:42:51 -04:00
# include "signal.h"
extern const unsigned long sigreturn_codes [ 17 ] ;
2005-04-16 15:20:36 -07:00
2013-07-24 00:29:18 +01:00
static unsigned long signal_return_offset ;
2005-04-16 15:20:36 -07:00
# ifdef CONFIG_IWMMXT
2017-06-30 18:56:09 +01:00
static int preserve_iwmmxt_context ( struct iwmmxt_sigframe __user * frame )
2005-04-16 15:20:36 -07:00
{
[PATCH] mm: arm ready for split ptlock
Prepare arm for the split page_table_lock: three issues.
Signal handling's preserve and restore of iwmmxt context currently involves
reading and writing that context to and from user space, while holding
page_table_lock to secure the user page(s) against kswapd. If we split the
lock, then the structure might span two pages, secured by to read into and
write from a kernel stack buffer, copying that out and in without locking (the
structure is 160 bytes in size, and here we're near the top of the kernel
stack). Or would the overhead be noticeable?
arm_syscall's cmpxchg emulation use pte_offset_map_lock, instead of
pte_offset_map and mm-wide page_table_lock; and strictly, it should now also
take mmap_sem before descending to pmd, to guard against another thread
munmapping, and the page table pulled out beneath this thread.
Updated two comments in fault-armv.c. adjust_pte is interesting, since its
modification of a pte in one part of the mm depends on the lock held when
calling update_mmu_cache for a pte in some other part of that mm. This can't
be done with a split page_table_lock (and we've already taken the lowest lock
in the hierarchy here): so we'll have to disable split on arm, unless
CONFIG_CPU_CACHE_VIPT to ensures adjust_pte never used.
Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-29 18:16:36 -07:00
char kbuf [ sizeof ( * frame ) + 8 ] ;
struct iwmmxt_sigframe * kframe ;
2017-06-30 18:56:59 +01:00
int err = 0 ;
2005-04-16 15:20:36 -07:00
/* the iWMMXt context must be 64 bit aligned */
[PATCH] mm: arm ready for split ptlock
Prepare arm for the split page_table_lock: three issues.
Signal handling's preserve and restore of iwmmxt context currently involves
reading and writing that context to and from user space, while holding
page_table_lock to secure the user page(s) against kswapd. If we split the
lock, then the structure might span two pages, secured by to read into and
write from a kernel stack buffer, copying that out and in without locking (the
structure is 160 bytes in size, and here we're near the top of the kernel
stack). Or would the overhead be noticeable?
arm_syscall's cmpxchg emulation use pte_offset_map_lock, instead of
pte_offset_map and mm-wide page_table_lock; and strictly, it should now also
take mmap_sem before descending to pmd, to guard against another thread
munmapping, and the page table pulled out beneath this thread.
Updated two comments in fault-armv.c. adjust_pte is interesting, since its
modification of a pte in one part of the mm depends on the lock held when
calling update_mmu_cache for a pte in some other part of that mm. This can't
be done with a split page_table_lock (and we've already taken the lowest lock
in the hierarchy here): so we'll have to disable split on arm, unless
CONFIG_CPU_CACHE_VIPT to ensures adjust_pte never used.
Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-29 18:16:36 -07:00
kframe = ( struct iwmmxt_sigframe * ) ( ( unsigned long ) ( kbuf + 8 ) & ~ 7 ) ;
2017-06-30 18:56:59 +01:00
if ( test_thread_flag ( TIF_USING_IWMMXT ) ) {
kframe - > magic = IWMMXT_MAGIC ;
kframe - > size = IWMMXT_STORAGE_SIZE ;
iwmmxt_task_copy ( current_thread_info ( ) , & kframe - > storage ) ;
} else {
/*
* For bug - compatibility with older kernels , some space
* has to be reserved for iWMMXt even if it ' s not used .
* Set the magic and size appropriately so that properly
* written userspace can skip it reliably :
*/
2018-09-11 10:12:05 +01:00
* kframe = ( struct iwmmxt_sigframe ) {
. magic = DUMMY_MAGIC ,
. size = IWMMXT_STORAGE_SIZE ,
} ;
2017-06-30 18:56:59 +01:00
}
2018-09-11 10:12:05 +01:00
err = __copy_to_user ( frame , kframe , sizeof ( * kframe ) ) ;
2017-06-30 18:56:59 +01:00
return err ;
2005-04-16 15:20:36 -07:00
}
2017-06-30 18:56:59 +01:00
static int restore_iwmmxt_context ( char __user * * auxp )
2005-04-16 15:20:36 -07:00
{
2017-06-30 18:56:59 +01:00
struct iwmmxt_sigframe __user * frame =
( struct iwmmxt_sigframe __user * ) * auxp ;
[PATCH] mm: arm ready for split ptlock
Prepare arm for the split page_table_lock: three issues.
Signal handling's preserve and restore of iwmmxt context currently involves
reading and writing that context to and from user space, while holding
page_table_lock to secure the user page(s) against kswapd. If we split the
lock, then the structure might span two pages, secured by to read into and
write from a kernel stack buffer, copying that out and in without locking (the
structure is 160 bytes in size, and here we're near the top of the kernel
stack). Or would the overhead be noticeable?
arm_syscall's cmpxchg emulation use pte_offset_map_lock, instead of
pte_offset_map and mm-wide page_table_lock; and strictly, it should now also
take mmap_sem before descending to pmd, to guard against another thread
munmapping, and the page table pulled out beneath this thread.
Updated two comments in fault-armv.c. adjust_pte is interesting, since its
modification of a pte in one part of the mm depends on the lock held when
calling update_mmu_cache for a pte in some other part of that mm. This can't
be done with a split page_table_lock (and we've already taken the lowest lock
in the hierarchy here): so we'll have to disable split on arm, unless
CONFIG_CPU_CACHE_VIPT to ensures adjust_pte never used.
Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-29 18:16:36 -07:00
char kbuf [ sizeof ( * frame ) + 8 ] ;
struct iwmmxt_sigframe * kframe ;
/* the iWMMXt context must be 64 bit aligned */
kframe = ( struct iwmmxt_sigframe * ) ( ( unsigned long ) ( kbuf + 8 ) & ~ 7 ) ;
if ( __copy_from_user ( kframe , frame , sizeof ( * frame ) ) )
return - 1 ;
2017-06-30 18:56:59 +01:00
/*
* For non - iWMMXt threads : a single iwmmxt_sigframe - sized dummy
* block is discarded for compatibility with setup_sigframe ( ) if
* present , but we don ' t mandate its presence . If some other
* magic is here , it ' s not for us :
*/
if ( ! test_thread_flag ( TIF_USING_IWMMXT ) & &
kframe - > magic ! = DUMMY_MAGIC )
return 0 ;
if ( kframe - > size ! = IWMMXT_STORAGE_SIZE )
[PATCH] mm: arm ready for split ptlock
Prepare arm for the split page_table_lock: three issues.
Signal handling's preserve and restore of iwmmxt context currently involves
reading and writing that context to and from user space, while holding
page_table_lock to secure the user page(s) against kswapd. If we split the
lock, then the structure might span two pages, secured by to read into and
write from a kernel stack buffer, copying that out and in without locking (the
structure is 160 bytes in size, and here we're near the top of the kernel
stack). Or would the overhead be noticeable?
arm_syscall's cmpxchg emulation use pte_offset_map_lock, instead of
pte_offset_map and mm-wide page_table_lock; and strictly, it should now also
take mmap_sem before descending to pmd, to guard against another thread
munmapping, and the page table pulled out beneath this thread.
Updated two comments in fault-armv.c. adjust_pte is interesting, since its
modification of a pte in one part of the mm depends on the lock held when
calling update_mmu_cache for a pte in some other part of that mm. This can't
be done with a split page_table_lock (and we've already taken the lowest lock
in the hierarchy here): so we'll have to disable split on arm, unless
CONFIG_CPU_CACHE_VIPT to ensures adjust_pte never used.
Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-29 18:16:36 -07:00
return - 1 ;
2017-06-30 18:56:59 +01:00
if ( test_thread_flag ( TIF_USING_IWMMXT ) ) {
if ( kframe - > magic ! = IWMMXT_MAGIC )
return - 1 ;
iwmmxt_task_restore ( current_thread_info ( ) , & kframe - > storage ) ;
}
* auxp + = IWMMXT_STORAGE_SIZE ;
[PATCH] mm: arm ready for split ptlock
Prepare arm for the split page_table_lock: three issues.
Signal handling's preserve and restore of iwmmxt context currently involves
reading and writing that context to and from user space, while holding
page_table_lock to secure the user page(s) against kswapd. If we split the
lock, then the structure might span two pages, secured by to read into and
write from a kernel stack buffer, copying that out and in without locking (the
structure is 160 bytes in size, and here we're near the top of the kernel
stack). Or would the overhead be noticeable?
arm_syscall's cmpxchg emulation use pte_offset_map_lock, instead of
pte_offset_map and mm-wide page_table_lock; and strictly, it should now also
take mmap_sem before descending to pmd, to guard against another thread
munmapping, and the page table pulled out beneath this thread.
Updated two comments in fault-armv.c. adjust_pte is interesting, since its
modification of a pte in one part of the mm depends on the lock held when
calling update_mmu_cache for a pte in some other part of that mm. This can't
be done with a split page_table_lock (and we've already taken the lowest lock
in the hierarchy here): so we'll have to disable split on arm, unless
CONFIG_CPU_CACHE_VIPT to ensures adjust_pte never used.
Signed-off-by: Hugh Dickins <hugh@veritas.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-29 18:16:36 -07:00
return 0 ;
2005-04-16 15:20:36 -07:00
}
# endif
2010-04-11 15:58:27 +01:00
# ifdef CONFIG_VFP
static int preserve_vfp_context ( struct vfp_sigframe __user * frame )
{
2018-09-11 10:12:18 +01:00
struct vfp_sigframe kframe ;
2010-04-11 15:58:27 +01:00
int err = 0 ;
2018-09-11 10:12:18 +01:00
memset ( & kframe , 0 , sizeof ( kframe ) ) ;
kframe . magic = VFP_MAGIC ;
kframe . size = VFP_STORAGE_SIZE ;
2010-04-11 15:58:27 +01:00
2018-09-11 10:12:18 +01:00
err = vfp_preserve_user_clear_hwstate ( & kframe . ufp , & kframe . ufp_exc ) ;
2012-04-23 15:38:28 +01:00
if ( err )
2018-09-11 10:12:18 +01:00
return err ;
2012-04-23 15:38:28 +01:00
2018-09-11 10:12:18 +01:00
return __copy_to_user ( frame , & kframe , sizeof ( kframe ) ) ;
2010-04-11 15:58:27 +01:00
}
2017-06-30 18:56:59 +01:00
static int restore_vfp_context ( char __user * * auxp )
2010-04-11 15:58:27 +01:00
{
2018-07-09 10:13:36 +01:00
struct vfp_sigframe frame ;
int err ;
2010-04-11 15:58:27 +01:00
2018-07-09 10:13:36 +01:00
err = __copy_from_user ( & frame , * auxp , sizeof ( frame ) ) ;
2010-04-11 15:58:27 +01:00
if ( err )
2018-07-09 10:13:36 +01:00
return err ;
if ( frame . magic ! = VFP_MAGIC | | frame . size ! = VFP_STORAGE_SIZE )
2010-04-11 15:58:27 +01:00
return - EINVAL ;
2018-07-09 10:13:36 +01:00
* auxp + = sizeof ( frame ) ;
return vfp_restore_user_hwstate ( & frame . ufp , & frame . ufp_exc ) ;
2010-04-11 15:58:27 +01:00
}
# endif
2005-04-16 15:20:36 -07:00
/*
* Do a signal return ; undo the signal stack . These are aligned to 64 - bit .
*/
2006-06-15 20:23:02 +01:00
static int restore_sigframe ( struct pt_regs * regs , struct sigframe __user * sf )
2005-04-16 15:20:36 -07:00
{
2018-07-09 10:05:22 +01:00
struct sigcontext context ;
2017-06-30 18:56:59 +01:00
char __user * aux ;
2006-06-15 20:23:02 +01:00
sigset_t set ;
int err ;
err = __copy_from_user ( & set , & sf - > uc . uc_sigmask , sizeof ( set ) ) ;
2012-04-27 13:58:59 -04:00
if ( err = = 0 )
2012-03-05 15:05:34 -08:00
set_current_blocked ( & set ) ;
2005-04-16 15:20:36 -07:00
2018-07-09 10:05:22 +01:00
err | = __copy_from_user ( & context , & sf - > uc . uc_mcontext , sizeof ( context ) ) ;
if ( err = = 0 ) {
regs - > ARM_r0 = context . arm_r0 ;
regs - > ARM_r1 = context . arm_r1 ;
regs - > ARM_r2 = context . arm_r2 ;
regs - > ARM_r3 = context . arm_r3 ;
regs - > ARM_r4 = context . arm_r4 ;
regs - > ARM_r5 = context . arm_r5 ;
regs - > ARM_r6 = context . arm_r6 ;
regs - > ARM_r7 = context . arm_r7 ;
regs - > ARM_r8 = context . arm_r8 ;
regs - > ARM_r9 = context . arm_r9 ;
regs - > ARM_r10 = context . arm_r10 ;
regs - > ARM_fp = context . arm_fp ;
regs - > ARM_ip = context . arm_ip ;
regs - > ARM_sp = context . arm_sp ;
regs - > ARM_lr = context . arm_lr ;
regs - > ARM_pc = context . arm_pc ;
regs - > ARM_cpsr = context . arm_cpsr ;
}
2005-04-16 15:20:36 -07:00
err | = ! valid_user_regs ( regs ) ;
2017-06-30 18:56:59 +01:00
aux = ( char __user * ) sf - > uc . uc_regspace ;
2005-04-16 15:20:36 -07:00
# ifdef CONFIG_IWMMXT
2017-06-30 18:56:59 +01:00
if ( err = = 0 )
err | = restore_iwmmxt_context ( & aux ) ;
2005-04-16 15:20:36 -07:00
# endif
# ifdef CONFIG_VFP
2010-04-11 15:58:27 +01:00
if ( err = = 0 )
2017-06-30 18:56:59 +01:00
err | = restore_vfp_context ( & aux ) ;
2005-04-16 15:20:36 -07:00
# endif
return err ;
}
asmlinkage int sys_sigreturn ( struct pt_regs * regs )
{
struct sigframe __user * frame ;
/* Always make any pending restarted system calls return -EINTR */
2015-02-12 15:01:14 -08:00
current - > restart_block . fn = do_no_restart_syscall ;
2005-04-16 15:20:36 -07:00
/*
* Since we stacked the signal on a 64 - bit boundary ,
* then ' sp ' should be word aligned here . If it ' s
* not , then the user is trying to mess with us .
*/
if ( regs - > ARM_sp & 7 )
goto badframe ;
frame = ( struct sigframe __user * ) regs - > ARM_sp ;
Remove 'type' argument from access_ok() function
Nobody has actually used the type (VERIFY_READ vs VERIFY_WRITE) argument
of the user address range verification function since we got rid of the
old racy i386-only code to walk page tables by hand.
It existed because the original 80386 would not honor the write protect
bit when in kernel mode, so you had to do COW by hand before doing any
user access. But we haven't supported that in a long time, and these
days the 'type' argument is a purely historical artifact.
A discussion about extending 'user_access_begin()' to do the range
checking resulted this patch, because there is no way we're going to
move the old VERIFY_xyz interface to that model. And it's best done at
the end of the merge window when I've done most of my merges, so let's
just get this done once and for all.
This patch was mostly done with a sed-script, with manual fix-ups for
the cases that weren't of the trivial 'access_ok(VERIFY_xyz' form.
There were a couple of notable cases:
- csky still had the old "verify_area()" name as an alias.
- the iter_iov code had magical hardcoded knowledge of the actual
values of VERIFY_{READ,WRITE} (not that they mattered, since nothing
really used it)
- microblaze used the type argument for a debug printout
but other than those oddities this should be a total no-op patch.
I tried to fix up all architectures, did fairly extensive grepping for
access_ok() uses, and the changes are trivial, but I may have missed
something. Any missed conversion should be trivially fixable, though.
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2019-01-03 18:57:57 -08:00
if ( ! access_ok ( frame , sizeof ( * frame ) ) )
2005-04-16 15:20:36 -07:00
goto badframe ;
2006-06-15 20:23:02 +01:00
if ( restore_sigframe ( regs , frame ) )
2005-04-16 15:20:36 -07:00
goto badframe ;
return regs - > ARM_r0 ;
badframe :
2019-05-23 10:17:27 -05:00
force_sig ( SIGSEGV ) ;
2005-04-16 15:20:36 -07:00
return 0 ;
}
asmlinkage int sys_rt_sigreturn ( struct pt_regs * regs )
{
struct rt_sigframe __user * frame ;
/* Always make any pending restarted system calls return -EINTR */
2015-02-12 15:01:14 -08:00
current - > restart_block . fn = do_no_restart_syscall ;
2005-04-16 15:20:36 -07:00
/*
* Since we stacked the signal on a 64 - bit boundary ,
* then ' sp ' should be word aligned here . If it ' s
* not , then the user is trying to mess with us .
*/
if ( regs - > ARM_sp & 7 )
goto badframe ;
frame = ( struct rt_sigframe __user * ) regs - > ARM_sp ;
Remove 'type' argument from access_ok() function
Nobody has actually used the type (VERIFY_READ vs VERIFY_WRITE) argument
of the user address range verification function since we got rid of the
old racy i386-only code to walk page tables by hand.
It existed because the original 80386 would not honor the write protect
bit when in kernel mode, so you had to do COW by hand before doing any
user access. But we haven't supported that in a long time, and these
days the 'type' argument is a purely historical artifact.
A discussion about extending 'user_access_begin()' to do the range
checking resulted this patch, because there is no way we're going to
move the old VERIFY_xyz interface to that model. And it's best done at
the end of the merge window when I've done most of my merges, so let's
just get this done once and for all.
This patch was mostly done with a sed-script, with manual fix-ups for
the cases that weren't of the trivial 'access_ok(VERIFY_xyz' form.
There were a couple of notable cases:
- csky still had the old "verify_area()" name as an alias.
- the iter_iov code had magical hardcoded knowledge of the actual
values of VERIFY_{READ,WRITE} (not that they mattered, since nothing
really used it)
- microblaze used the type argument for a debug printout
but other than those oddities this should be a total no-op patch.
I tried to fix up all architectures, did fairly extensive grepping for
access_ok() uses, and the changes are trivial, but I may have missed
something. Any missed conversion should be trivially fixable, though.
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2019-01-03 18:57:57 -08:00
if ( ! access_ok ( frame , sizeof ( * frame ) ) )
2005-04-16 15:20:36 -07:00
goto badframe ;
2006-06-15 20:23:02 +01:00
if ( restore_sigframe ( regs , & frame - > sig ) )
2005-04-16 15:20:36 -07:00
goto badframe ;
2012-12-23 01:52:54 -05:00
if ( restore_altstack ( & frame - > sig . uc . uc_stack ) )
2005-04-16 15:20:36 -07:00
goto badframe ;
return regs - > ARM_r0 ;
badframe :
2019-05-23 10:17:27 -05:00
force_sig ( SIGSEGV ) ;
2005-04-16 15:20:36 -07:00
return 0 ;
}
static int
2006-06-15 20:28:03 +01:00
setup_sigframe ( struct sigframe __user * sf , struct pt_regs * regs , sigset_t * set )
2005-04-16 15:20:36 -07:00
{
2006-06-24 23:46:21 +01:00
struct aux_sigframe __user * aux ;
2018-09-11 10:11:06 +01:00
struct sigcontext context ;
2005-04-16 15:20:36 -07:00
int err = 0 ;
2018-09-11 10:11:06 +01:00
context = ( struct sigcontext ) {
. arm_r0 = regs - > ARM_r0 ,
. arm_r1 = regs - > ARM_r1 ,
. arm_r2 = regs - > ARM_r2 ,
. arm_r3 = regs - > ARM_r3 ,
. arm_r4 = regs - > ARM_r4 ,
. arm_r5 = regs - > ARM_r5 ,
. arm_r6 = regs - > ARM_r6 ,
. arm_r7 = regs - > ARM_r7 ,
. arm_r8 = regs - > ARM_r8 ,
. arm_r9 = regs - > ARM_r9 ,
. arm_r10 = regs - > ARM_r10 ,
. arm_fp = regs - > ARM_fp ,
. arm_ip = regs - > ARM_ip ,
. arm_sp = regs - > ARM_sp ,
. arm_lr = regs - > ARM_lr ,
. arm_pc = regs - > ARM_pc ,
. arm_cpsr = regs - > ARM_cpsr ,
. trap_no = current - > thread . trap_no ,
. error_code = current - > thread . error_code ,
. fault_address = current - > thread . address ,
. oldmask = set - > sig [ 0 ] ,
} ;
err | = __copy_to_user ( & sf - > uc . uc_mcontext , & context , sizeof ( context ) ) ;
2006-06-15 20:28:03 +01:00
err | = __copy_to_user ( & sf - > uc . uc_sigmask , set , sizeof ( * set ) ) ;
2005-04-16 15:20:36 -07:00
2006-06-24 23:46:21 +01:00
aux = ( struct aux_sigframe __user * ) sf - > uc . uc_regspace ;
2005-04-16 15:20:36 -07:00
# ifdef CONFIG_IWMMXT
2017-06-30 18:56:59 +01:00
if ( err = = 0 )
2005-04-16 15:20:36 -07:00
err | = preserve_iwmmxt_context ( & aux - > iwmmxt ) ;
# endif
# ifdef CONFIG_VFP
2010-04-11 15:58:27 +01:00
if ( err = = 0 )
err | = preserve_vfp_context ( & aux - > vfp ) ;
2005-04-16 15:20:36 -07:00
# endif
2018-09-11 10:13:11 +01:00
err | = __put_user ( 0 , & aux - > end_magic ) ;
2005-04-16 15:20:36 -07:00
return err ;
}
static inline void __user *
2012-11-07 17:53:13 -05:00
get_sigframe ( struct ksignal * ksig , struct pt_regs * regs , int framesize )
2005-04-16 15:20:36 -07:00
{
2012-11-07 17:53:13 -05:00
unsigned long sp = sigsp ( regs - > ARM_sp , ksig ) ;
2005-04-16 15:20:36 -07:00
void __user * frame ;
/*
* ATPCS B01 mandates 8 - byte alignment
*/
frame = ( void __user * ) ( ( sp - framesize ) & ~ 7 ) ;
/*
* Check that we can actually write to the signal frame .
*/
Remove 'type' argument from access_ok() function
Nobody has actually used the type (VERIFY_READ vs VERIFY_WRITE) argument
of the user address range verification function since we got rid of the
old racy i386-only code to walk page tables by hand.
It existed because the original 80386 would not honor the write protect
bit when in kernel mode, so you had to do COW by hand before doing any
user access. But we haven't supported that in a long time, and these
days the 'type' argument is a purely historical artifact.
A discussion about extending 'user_access_begin()' to do the range
checking resulted this patch, because there is no way we're going to
move the old VERIFY_xyz interface to that model. And it's best done at
the end of the merge window when I've done most of my merges, so let's
just get this done once and for all.
This patch was mostly done with a sed-script, with manual fix-ups for
the cases that weren't of the trivial 'access_ok(VERIFY_xyz' form.
There were a couple of notable cases:
- csky still had the old "verify_area()" name as an alias.
- the iter_iov code had magical hardcoded knowledge of the actual
values of VERIFY_{READ,WRITE} (not that they mattered, since nothing
really used it)
- microblaze used the type argument for a debug printout
but other than those oddities this should be a total no-op patch.
I tried to fix up all architectures, did fairly extensive grepping for
access_ok() uses, and the changes are trivial, but I may have missed
something. Any missed conversion should be trivially fixable, though.
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2019-01-03 18:57:57 -08:00
if ( ! access_ok ( frame , framesize ) )
2005-04-16 15:20:36 -07:00
frame = NULL ;
return frame ;
}
static int
2012-11-07 17:53:13 -05:00
setup_return ( struct pt_regs * regs , struct ksignal * ksig ,
unsigned long __user * rc , void __user * frame )
2005-04-16 15:20:36 -07:00
{
2012-11-07 17:53:13 -05:00
unsigned long handler = ( unsigned long ) ksig - > ka . sa . sa_handler ;
2017-08-09 23:42:51 -04:00
unsigned long handler_fdpic_GOT = 0 ;
2005-04-16 15:20:36 -07:00
unsigned long retcode ;
2017-08-09 23:42:51 -04:00
unsigned int idx , thumb = 0 ;
2011-02-20 12:22:52 +00:00
unsigned long cpsr = regs - > ARM_cpsr & ~ ( PSR_f | PSR_E_BIT ) ;
2017-08-09 23:42:51 -04:00
bool fdpic = IS_ENABLED ( CONFIG_BINFMT_ELF_FDPIC ) & &
( current - > personality & FDPIC_FUNCPTRS ) ;
if ( fdpic ) {
unsigned long __user * fdpic_func_desc =
( unsigned long __user * ) handler ;
if ( __get_user ( handler , & fdpic_func_desc [ 0 ] ) | |
__get_user ( handler_fdpic_GOT , & fdpic_func_desc [ 1 ] ) )
return 1 ;
}
2011-02-20 12:22:52 +00:00
cpsr | = PSR_ENDSTATE ;
2005-04-16 15:20:36 -07:00
/*
* Maybe we need to deliver a 32 - bit signal to a 26 - bit task .
*/
2012-11-07 17:53:13 -05:00
if ( ksig - > ka . sa . sa_flags & SA_THIRTYTWO )
2005-04-16 15:20:36 -07:00
cpsr = ( cpsr & ~ MODE_MASK ) | USR_MODE ;
# ifdef CONFIG_ARM_THUMB
if ( elf_hwcap & HWCAP_THUMB ) {
/*
* The LSB of the handler determines if we ' re going to
* be using THUMB or ARM mode for this signal handler .
*/
thumb = handler & 1 ;
2013-11-06 18:38:05 +01:00
/*
2015-09-11 16:44:02 +01:00
* Clear the If - Then Thumb - 2 execution state . ARM spec
* requires this to be all 000 s in ARM mode . Snapdragon
* S4 / Krait misbehaves on a Thumb = > ARM signal transition
* without this .
*
* We must do this whenever we are running on a Thumb - 2
* capable CPU , which includes ARMv6T2 . However , we elect
2015-09-16 11:08:49 +01:00
* to always do this to simplify the code ; this field is
* marked UNK / SBZP for older architectures .
2013-11-06 18:38:05 +01:00
*/
cpsr & = ~ PSR_IT_MASK ;
if ( thumb ) {
cpsr | = PSR_T_BIT ;
2009-05-30 14:00:15 +01:00
} else
2005-04-16 15:20:36 -07:00
cpsr & = ~ PSR_T_BIT ;
}
# endif
2012-11-07 17:53:13 -05:00
if ( ksig - > ka . sa . sa_flags & SA_RESTORER ) {
retcode = ( unsigned long ) ksig - > ka . sa . sa_restorer ;
2017-08-09 23:42:51 -04:00
if ( fdpic ) {
/*
* We need code to load the function descriptor .
* That code follows the standard sigreturn code
* ( 6 words ) , and is made of 3 + 2 words for each
* variant . The 4 th copied word is the actual FD
* address that the assembly code expects .
*/
idx = 6 + thumb * 3 ;
if ( ksig - > ka . sa . sa_flags & SA_SIGINFO )
idx + = 5 ;
if ( __put_user ( sigreturn_codes [ idx ] , rc ) | |
__put_user ( sigreturn_codes [ idx + 1 ] , rc + 1 ) | |
__put_user ( sigreturn_codes [ idx + 2 ] , rc + 2 ) | |
__put_user ( retcode , rc + 3 ) )
return 1 ;
goto rc_finish ;
}
2005-04-16 15:20:36 -07:00
} else {
2017-08-09 23:42:51 -04:00
idx = thumb < < 1 ;
2012-11-07 17:53:13 -05:00
if ( ksig - > ka . sa . sa_flags & SA_SIGINFO )
2006-01-18 22:38:47 +00:00
idx + = 3 ;
2005-04-16 15:20:36 -07:00
2013-04-18 18:37:24 +01:00
/*
* Put the sigreturn code on the stack no matter which return
* mechanism we use in order to remain ABI compliant
*/
2006-01-18 22:38:47 +00:00
if ( __put_user ( sigreturn_codes [ idx ] , rc ) | |
__put_user ( sigreturn_codes [ idx + 1 ] , rc + 1 ) )
2005-04-16 15:20:36 -07:00
return 1 ;
2017-08-09 23:42:51 -04:00
rc_finish :
2013-08-03 10:39:51 +01:00
# ifdef CONFIG_MMU
if ( cpsr & MODE32_BIT ) {
2013-07-24 00:29:18 +01:00
struct mm_struct * mm = current - > mm ;
2005-06-22 20:26:05 +01:00
/*
2013-07-24 00:29:18 +01:00
* 32 - bit code can use the signal return page
* except when the MPU has protected the vectors
* page from PL0
2005-06-22 20:26:05 +01:00
*/
2013-07-24 00:29:18 +01:00
retcode = mm - > context . sigpage + signal_return_offset +
( idx < < 2 ) + thumb ;
2013-08-03 10:39:51 +01:00
} else
# endif
{
2005-06-22 20:26:05 +01:00
/*
* Ensure that the instruction cache sees
* the return code written onto the stack .
*/
flush_icache_range ( ( unsigned long ) rc ,
2017-08-09 23:42:51 -04:00
( unsigned long ) ( rc + 3 ) ) ;
2005-06-22 20:26:05 +01:00
retcode = ( ( unsigned long ) rc ) + thumb ;
}
2005-04-16 15:20:36 -07:00
}
2014-07-13 15:24:03 +02:00
regs - > ARM_r0 = ksig - > sig ;
2005-04-16 15:20:36 -07:00
regs - > ARM_sp = ( unsigned long ) frame ;
regs - > ARM_lr = retcode ;
regs - > ARM_pc = handler ;
2017-08-09 23:42:51 -04:00
if ( fdpic )
regs - > ARM_r9 = handler_fdpic_GOT ;
2005-04-16 15:20:36 -07:00
regs - > ARM_cpsr = cpsr ;
return 0 ;
}
static int
2012-11-07 17:53:13 -05:00
setup_frame ( struct ksignal * ksig , sigset_t * set , struct pt_regs * regs )
2005-04-16 15:20:36 -07:00
{
2012-11-07 17:53:13 -05:00
struct sigframe __user * frame = get_sigframe ( ksig , regs , sizeof ( * frame ) ) ;
2005-04-16 15:20:36 -07:00
int err = 0 ;
if ( ! frame )
return 1 ;
2006-06-24 22:41:09 +01:00
/*
* Set uc . uc_flags to a value which sc . trap_no would never have .
*/
2018-09-11 10:13:11 +01:00
err = __put_user ( 0x5ac3c35a , & frame - > uc . uc_flags ) ;
2005-04-16 15:20:36 -07:00
2006-06-15 20:28:03 +01:00
err | = setup_sigframe ( frame , regs , set ) ;
2005-04-16 15:20:36 -07:00
if ( err = = 0 )
2012-11-07 17:53:13 -05:00
err = setup_return ( regs , ksig , frame - > retcode , frame ) ;
2005-04-16 15:20:36 -07:00
return err ;
}
static int
2012-11-07 17:53:13 -05:00
setup_rt_frame ( struct ksignal * ksig , sigset_t * set , struct pt_regs * regs )
2005-04-16 15:20:36 -07:00
{
2012-11-07 17:53:13 -05:00
struct rt_sigframe __user * frame = get_sigframe ( ksig , regs , sizeof ( * frame ) ) ;
2005-04-16 15:20:36 -07:00
int err = 0 ;
if ( ! frame )
return 1 ;
2012-11-07 17:53:13 -05:00
err | = copy_siginfo_to_user ( & frame - > info , & ksig - > info ) ;
2005-04-16 15:20:36 -07:00
2018-09-11 10:13:11 +01:00
err | = __put_user ( 0 , & frame - > sig . uc . uc_flags ) ;
err | = __put_user ( NULL , & frame - > sig . uc . uc_link ) ;
2005-04-16 15:20:36 -07:00
2012-12-23 01:52:54 -05:00
err | = __save_altstack ( & frame - > sig . uc . uc_stack , regs - > ARM_sp ) ;
2006-06-15 20:28:03 +01:00
err | = setup_sigframe ( & frame - > sig , regs , set ) ;
2005-04-16 15:20:36 -07:00
if ( err = = 0 )
2012-11-07 17:53:13 -05:00
err = setup_return ( regs , ksig , frame - > sig . retcode , frame ) ;
2005-04-16 15:20:36 -07:00
if ( err = = 0 ) {
/*
* For realtime signals we must also set the second and third
* arguments for the signal handler .
* - - Peter Maydell < pmaydell @ chiark . greenend . org . uk > 2000 - 12 - 06
*/
regs - > ARM_r1 = ( unsigned long ) & frame - > info ;
2006-06-15 20:18:25 +01:00
regs - > ARM_r2 = ( unsigned long ) & frame - > sig . uc ;
2005-04-16 15:20:36 -07:00
}
return err ;
}
/*
* OK , we ' re invoking a handler
*/
2012-11-07 17:53:13 -05:00
static void handle_signal ( struct ksignal * ksig , struct pt_regs * regs )
2005-04-16 15:20:36 -07:00
{
2012-05-02 09:59:21 -04:00
sigset_t * oldset = sigmask_to_save ( ) ;
2005-04-16 15:20:36 -07:00
int ret ;
2018-06-02 08:43:55 -04:00
/*
2019-03-05 14:47:53 -05:00
* Perform fixup for the pre - signal frame .
2018-06-02 08:43:55 -04:00
*/
2018-06-22 11:45:07 +01:00
rseq_signal_deliver ( ksig , regs ) ;
2018-06-02 08:43:55 -04:00
2005-04-16 15:20:36 -07:00
/*
* Set up the stack frame
*/
2012-11-07 17:53:13 -05:00
if ( ksig - > ka . sa . sa_flags & SA_SIGINFO )
ret = setup_rt_frame ( ksig , oldset , regs ) ;
2005-04-16 15:20:36 -07:00
else
2012-11-07 17:53:13 -05:00
ret = setup_frame ( ksig , oldset , regs ) ;
2005-04-16 15:20:36 -07:00
/*
* Check that the resulting registers are actually sane .
*/
ret | = ! valid_user_regs ( regs ) ;
2012-11-07 17:53:13 -05:00
signal_setup_done ( ret , ksig , 0 ) ;
2005-04-16 15:20:36 -07:00
}
/*
* Note that ' init ' is a special process : it doesn ' t get signals it doesn ' t
* want to handle . Thus you cannot kill init even with a SIGKILL even by
* mistake .
*
* Note that we go through the signals twice : once to check the signals that
* the kernel can handle , and then we build all the user - level signal handling
* stack - frames in one go after that .
*/
2012-07-19 17:48:21 +01:00
static int do_signal ( struct pt_regs * regs , int syscall )
2005-04-16 15:20:36 -07:00
{
ARM: 6892/1: handle ptrace requests to change PC during interrupted system calls
GDB's interrupt.exp test cases currenly fail on ARM. The problem is how do_signal
handled restarting interrupted system calls:
The entry.S assembler code determines that we come from a system call; and that
information is passed as "syscall" parameter to do_signal. That routine then
calls get_signal_to_deliver [*] and if a signal is to be delivered, calls into
handle_signal. If a system call is to be restarted either after the signal
handler returns, or if no handler is to be called in the first place, the PC
is updated after the get_signal_to_deliver call, either in handle_signal (if
we have a handler) or at the end of do_signal (otherwise).
Now the problem is that during [*], the call to get_signal_to_deliver, a ptrace
intercept may happen. During this intercept, the debugger may change registers,
including the PC. This is done by GDB if it wants to execute an "inferior call",
i.e. the execution of some code in the debugged program triggered by GDB.
To this purpose, GDB will save all registers, allocate a stack frame, set up
PC and arguments as appropriate for the call, and point the link register to
a dummy breakpoint instruction. Once the process is restarted, it will execute
the call and then trap back to the debugger, at which point GDB will restore
all registers and continue original execution.
This generally works fine. However, now consider what happens when GDB attempts
to do exactly that while the process was interrupted during execution of a to-be-
restarted system call: do_signal is called with the syscall flag set; it calls
get_signal_to_deliver, at which point the debugger takes over and changes the PC
to point to a completely different place. Now get_signal_to_deliver returns
without a signal to deliver; but now do_signal decides it should be restarting
a system call, and decrements the PC by 2 or 4 -- so it now points to 2 or 4
bytes before the function GDB wants to call -- which leads to a subsequent crash.
To fix this problem, two things need to be supported:
- do_signal must be able to recognize that get_signal_to_deliver changed the PC
to a different location, and skip the restart-syscall sequence
- once the debugger has restored all registers at the end of the inferior call
sequence, do_signal must recognize that *now* it needs to restart the pending
system call, even though it was now entered from a breakpoint instead of an
actual svc instruction
This set of issues is solved on other platforms, usually by one of two
mechanisms:
- The status information "do_signal is handling a system call that may need
restarting" is itself carried in some register that can be accessed via
ptrace. This is e.g. on Intel the "orig_eax" register; on Sparc the kernel
defines a magic extra bit in the flags register for this purpose.
This allows GDB to manage that state: reset it when doing an inferior call,
and restore it after the call is finished.
- On s390, do_signal transparently handles this problem without requiring
GDB interaction, by performing system call restarting in the following
way: first, adjust the PC as necessary for restarting the call. Then,
call get_signal_to_deliver; and finally just continue execution at the
PC. This way, if GDB does not change the PC, everything is as before.
If GDB *does* change the PC, execution will simply continue there --
and once GDB restores the PC it saved at that point, it will automatically
point to the *restarted* system call. (There is the minor twist how to
handle system calls that do *not* need restarting -- do_signal will undo
the PC change in this case, after get_signal_to_deliver has returned, and
only if ptrace did not change the PC during that call.)
Because there does not appear to be any obvious register to carry the
syscall-restart information on ARM, we'd either have to introduce a new
artificial ptrace register just for that purpose, or else handle the issue
transparently like on s390. The patch below implements the second option;
using this patch makes the interrupt.exp test cases pass on ARM, with no
regression in the GDB test suite otherwise.
Cc: patches@linaro.org
Signed-off-by: Ulrich Weigand <ulrich.weigand@linaro.org>
Signed-off-by: Arnd Bergmann <arnd.bergmann@linaro.org>
Signed-off-by: Russell King <rmk+kernel@arm.linux.org.uk>
2011-05-03 18:32:55 +01:00
unsigned int retval = 0 , continue_addr = 0 , restart_addr = 0 ;
2012-11-07 17:53:13 -05:00
struct ksignal ksig ;
2012-07-19 17:48:21 +01:00
int restart = 0 ;
2005-04-16 15:20:36 -07:00
ARM: 6892/1: handle ptrace requests to change PC during interrupted system calls
GDB's interrupt.exp test cases currenly fail on ARM. The problem is how do_signal
handled restarting interrupted system calls:
The entry.S assembler code determines that we come from a system call; and that
information is passed as "syscall" parameter to do_signal. That routine then
calls get_signal_to_deliver [*] and if a signal is to be delivered, calls into
handle_signal. If a system call is to be restarted either after the signal
handler returns, or if no handler is to be called in the first place, the PC
is updated after the get_signal_to_deliver call, either in handle_signal (if
we have a handler) or at the end of do_signal (otherwise).
Now the problem is that during [*], the call to get_signal_to_deliver, a ptrace
intercept may happen. During this intercept, the debugger may change registers,
including the PC. This is done by GDB if it wants to execute an "inferior call",
i.e. the execution of some code in the debugged program triggered by GDB.
To this purpose, GDB will save all registers, allocate a stack frame, set up
PC and arguments as appropriate for the call, and point the link register to
a dummy breakpoint instruction. Once the process is restarted, it will execute
the call and then trap back to the debugger, at which point GDB will restore
all registers and continue original execution.
This generally works fine. However, now consider what happens when GDB attempts
to do exactly that while the process was interrupted during execution of a to-be-
restarted system call: do_signal is called with the syscall flag set; it calls
get_signal_to_deliver, at which point the debugger takes over and changes the PC
to point to a completely different place. Now get_signal_to_deliver returns
without a signal to deliver; but now do_signal decides it should be restarting
a system call, and decrements the PC by 2 or 4 -- so it now points to 2 or 4
bytes before the function GDB wants to call -- which leads to a subsequent crash.
To fix this problem, two things need to be supported:
- do_signal must be able to recognize that get_signal_to_deliver changed the PC
to a different location, and skip the restart-syscall sequence
- once the debugger has restored all registers at the end of the inferior call
sequence, do_signal must recognize that *now* it needs to restart the pending
system call, even though it was now entered from a breakpoint instead of an
actual svc instruction
This set of issues is solved on other platforms, usually by one of two
mechanisms:
- The status information "do_signal is handling a system call that may need
restarting" is itself carried in some register that can be accessed via
ptrace. This is e.g. on Intel the "orig_eax" register; on Sparc the kernel
defines a magic extra bit in the flags register for this purpose.
This allows GDB to manage that state: reset it when doing an inferior call,
and restore it after the call is finished.
- On s390, do_signal transparently handles this problem without requiring
GDB interaction, by performing system call restarting in the following
way: first, adjust the PC as necessary for restarting the call. Then,
call get_signal_to_deliver; and finally just continue execution at the
PC. This way, if GDB does not change the PC, everything is as before.
If GDB *does* change the PC, execution will simply continue there --
and once GDB restores the PC it saved at that point, it will automatically
point to the *restarted* system call. (There is the minor twist how to
handle system calls that do *not* need restarting -- do_signal will undo
the PC change in this case, after get_signal_to_deliver has returned, and
only if ptrace did not change the PC during that call.)
Because there does not appear to be any obvious register to carry the
syscall-restart information on ARM, we'd either have to introduce a new
artificial ptrace register just for that purpose, or else handle the issue
transparently like on s390. The patch below implements the second option;
using this patch makes the interrupt.exp test cases pass on ARM, with no
regression in the GDB test suite otherwise.
Cc: patches@linaro.org
Signed-off-by: Ulrich Weigand <ulrich.weigand@linaro.org>
Signed-off-by: Arnd Bergmann <arnd.bergmann@linaro.org>
Signed-off-by: Russell King <rmk+kernel@arm.linux.org.uk>
2011-05-03 18:32:55 +01:00
/*
* If we were from a system call , check for system call restarting . . .
*/
if ( syscall ) {
continue_addr = regs - > ARM_pc ;
restart_addr = continue_addr - ( thumb_mode ( regs ) ? 2 : 4 ) ;
retval = regs - > ARM_r0 ;
/*
* Prepare for system call restart . We do this here so that a
* debugger will see the already changed PSW .
*/
switch ( retval ) {
2012-07-19 17:48:21 +01:00
case - ERESTART_RESTARTBLOCK :
2012-07-19 17:48:50 +01:00
restart - = 2 ;
2020-08-23 17:36:59 -05:00
fallthrough ;
ARM: 6892/1: handle ptrace requests to change PC during interrupted system calls
GDB's interrupt.exp test cases currenly fail on ARM. The problem is how do_signal
handled restarting interrupted system calls:
The entry.S assembler code determines that we come from a system call; and that
information is passed as "syscall" parameter to do_signal. That routine then
calls get_signal_to_deliver [*] and if a signal is to be delivered, calls into
handle_signal. If a system call is to be restarted either after the signal
handler returns, or if no handler is to be called in the first place, the PC
is updated after the get_signal_to_deliver call, either in handle_signal (if
we have a handler) or at the end of do_signal (otherwise).
Now the problem is that during [*], the call to get_signal_to_deliver, a ptrace
intercept may happen. During this intercept, the debugger may change registers,
including the PC. This is done by GDB if it wants to execute an "inferior call",
i.e. the execution of some code in the debugged program triggered by GDB.
To this purpose, GDB will save all registers, allocate a stack frame, set up
PC and arguments as appropriate for the call, and point the link register to
a dummy breakpoint instruction. Once the process is restarted, it will execute
the call and then trap back to the debugger, at which point GDB will restore
all registers and continue original execution.
This generally works fine. However, now consider what happens when GDB attempts
to do exactly that while the process was interrupted during execution of a to-be-
restarted system call: do_signal is called with the syscall flag set; it calls
get_signal_to_deliver, at which point the debugger takes over and changes the PC
to point to a completely different place. Now get_signal_to_deliver returns
without a signal to deliver; but now do_signal decides it should be restarting
a system call, and decrements the PC by 2 or 4 -- so it now points to 2 or 4
bytes before the function GDB wants to call -- which leads to a subsequent crash.
To fix this problem, two things need to be supported:
- do_signal must be able to recognize that get_signal_to_deliver changed the PC
to a different location, and skip the restart-syscall sequence
- once the debugger has restored all registers at the end of the inferior call
sequence, do_signal must recognize that *now* it needs to restart the pending
system call, even though it was now entered from a breakpoint instead of an
actual svc instruction
This set of issues is solved on other platforms, usually by one of two
mechanisms:
- The status information "do_signal is handling a system call that may need
restarting" is itself carried in some register that can be accessed via
ptrace. This is e.g. on Intel the "orig_eax" register; on Sparc the kernel
defines a magic extra bit in the flags register for this purpose.
This allows GDB to manage that state: reset it when doing an inferior call,
and restore it after the call is finished.
- On s390, do_signal transparently handles this problem without requiring
GDB interaction, by performing system call restarting in the following
way: first, adjust the PC as necessary for restarting the call. Then,
call get_signal_to_deliver; and finally just continue execution at the
PC. This way, if GDB does not change the PC, everything is as before.
If GDB *does* change the PC, execution will simply continue there --
and once GDB restores the PC it saved at that point, it will automatically
point to the *restarted* system call. (There is the minor twist how to
handle system calls that do *not* need restarting -- do_signal will undo
the PC change in this case, after get_signal_to_deliver has returned, and
only if ptrace did not change the PC during that call.)
Because there does not appear to be any obvious register to carry the
syscall-restart information on ARM, we'd either have to introduce a new
artificial ptrace register just for that purpose, or else handle the issue
transparently like on s390. The patch below implements the second option;
using this patch makes the interrupt.exp test cases pass on ARM, with no
regression in the GDB test suite otherwise.
Cc: patches@linaro.org
Signed-off-by: Ulrich Weigand <ulrich.weigand@linaro.org>
Signed-off-by: Arnd Bergmann <arnd.bergmann@linaro.org>
Signed-off-by: Russell King <rmk+kernel@arm.linux.org.uk>
2011-05-03 18:32:55 +01:00
case - ERESTARTNOHAND :
case - ERESTARTSYS :
case - ERESTARTNOINTR :
2012-07-19 17:48:21 +01:00
restart + + ;
ARM: 6892/1: handle ptrace requests to change PC during interrupted system calls
GDB's interrupt.exp test cases currenly fail on ARM. The problem is how do_signal
handled restarting interrupted system calls:
The entry.S assembler code determines that we come from a system call; and that
information is passed as "syscall" parameter to do_signal. That routine then
calls get_signal_to_deliver [*] and if a signal is to be delivered, calls into
handle_signal. If a system call is to be restarted either after the signal
handler returns, or if no handler is to be called in the first place, the PC
is updated after the get_signal_to_deliver call, either in handle_signal (if
we have a handler) or at the end of do_signal (otherwise).
Now the problem is that during [*], the call to get_signal_to_deliver, a ptrace
intercept may happen. During this intercept, the debugger may change registers,
including the PC. This is done by GDB if it wants to execute an "inferior call",
i.e. the execution of some code in the debugged program triggered by GDB.
To this purpose, GDB will save all registers, allocate a stack frame, set up
PC and arguments as appropriate for the call, and point the link register to
a dummy breakpoint instruction. Once the process is restarted, it will execute
the call and then trap back to the debugger, at which point GDB will restore
all registers and continue original execution.
This generally works fine. However, now consider what happens when GDB attempts
to do exactly that while the process was interrupted during execution of a to-be-
restarted system call: do_signal is called with the syscall flag set; it calls
get_signal_to_deliver, at which point the debugger takes over and changes the PC
to point to a completely different place. Now get_signal_to_deliver returns
without a signal to deliver; but now do_signal decides it should be restarting
a system call, and decrements the PC by 2 or 4 -- so it now points to 2 or 4
bytes before the function GDB wants to call -- which leads to a subsequent crash.
To fix this problem, two things need to be supported:
- do_signal must be able to recognize that get_signal_to_deliver changed the PC
to a different location, and skip the restart-syscall sequence
- once the debugger has restored all registers at the end of the inferior call
sequence, do_signal must recognize that *now* it needs to restart the pending
system call, even though it was now entered from a breakpoint instead of an
actual svc instruction
This set of issues is solved on other platforms, usually by one of two
mechanisms:
- The status information "do_signal is handling a system call that may need
restarting" is itself carried in some register that can be accessed via
ptrace. This is e.g. on Intel the "orig_eax" register; on Sparc the kernel
defines a magic extra bit in the flags register for this purpose.
This allows GDB to manage that state: reset it when doing an inferior call,
and restore it after the call is finished.
- On s390, do_signal transparently handles this problem without requiring
GDB interaction, by performing system call restarting in the following
way: first, adjust the PC as necessary for restarting the call. Then,
call get_signal_to_deliver; and finally just continue execution at the
PC. This way, if GDB does not change the PC, everything is as before.
If GDB *does* change the PC, execution will simply continue there --
and once GDB restores the PC it saved at that point, it will automatically
point to the *restarted* system call. (There is the minor twist how to
handle system calls that do *not* need restarting -- do_signal will undo
the PC change in this case, after get_signal_to_deliver has returned, and
only if ptrace did not change the PC during that call.)
Because there does not appear to be any obvious register to carry the
syscall-restart information on ARM, we'd either have to introduce a new
artificial ptrace register just for that purpose, or else handle the issue
transparently like on s390. The patch below implements the second option;
using this patch makes the interrupt.exp test cases pass on ARM, with no
regression in the GDB test suite otherwise.
Cc: patches@linaro.org
Signed-off-by: Ulrich Weigand <ulrich.weigand@linaro.org>
Signed-off-by: Arnd Bergmann <arnd.bergmann@linaro.org>
Signed-off-by: Russell King <rmk+kernel@arm.linux.org.uk>
2011-05-03 18:32:55 +01:00
regs - > ARM_r0 = regs - > ARM_ORIG_r0 ;
regs - > ARM_pc = restart_addr ;
break ;
}
}
/*
* Get the signal to deliver . When running under ptrace , at this
* point the debugger may change all our registers . . .
*/
2012-07-19 17:48:21 +01:00
/*
* Depending on the signal settings we may need to revert the
* decision to restart the system call . But skip this if a
* debugger has chosen to restart at a different PC .
*/
2012-11-07 17:53:13 -05:00
if ( get_signal ( & ksig ) ) {
/* handler */
if ( unlikely ( restart ) & & regs - > ARM_pc = = restart_addr ) {
2012-07-19 17:46:44 +01:00
if ( retval = = - ERESTARTNOHAND | |
retval = = - ERESTART_RESTARTBLOCK
ARM: 6892/1: handle ptrace requests to change PC during interrupted system calls
GDB's interrupt.exp test cases currenly fail on ARM. The problem is how do_signal
handled restarting interrupted system calls:
The entry.S assembler code determines that we come from a system call; and that
information is passed as "syscall" parameter to do_signal. That routine then
calls get_signal_to_deliver [*] and if a signal is to be delivered, calls into
handle_signal. If a system call is to be restarted either after the signal
handler returns, or if no handler is to be called in the first place, the PC
is updated after the get_signal_to_deliver call, either in handle_signal (if
we have a handler) or at the end of do_signal (otherwise).
Now the problem is that during [*], the call to get_signal_to_deliver, a ptrace
intercept may happen. During this intercept, the debugger may change registers,
including the PC. This is done by GDB if it wants to execute an "inferior call",
i.e. the execution of some code in the debugged program triggered by GDB.
To this purpose, GDB will save all registers, allocate a stack frame, set up
PC and arguments as appropriate for the call, and point the link register to
a dummy breakpoint instruction. Once the process is restarted, it will execute
the call and then trap back to the debugger, at which point GDB will restore
all registers and continue original execution.
This generally works fine. However, now consider what happens when GDB attempts
to do exactly that while the process was interrupted during execution of a to-be-
restarted system call: do_signal is called with the syscall flag set; it calls
get_signal_to_deliver, at which point the debugger takes over and changes the PC
to point to a completely different place. Now get_signal_to_deliver returns
without a signal to deliver; but now do_signal decides it should be restarting
a system call, and decrements the PC by 2 or 4 -- so it now points to 2 or 4
bytes before the function GDB wants to call -- which leads to a subsequent crash.
To fix this problem, two things need to be supported:
- do_signal must be able to recognize that get_signal_to_deliver changed the PC
to a different location, and skip the restart-syscall sequence
- once the debugger has restored all registers at the end of the inferior call
sequence, do_signal must recognize that *now* it needs to restart the pending
system call, even though it was now entered from a breakpoint instead of an
actual svc instruction
This set of issues is solved on other platforms, usually by one of two
mechanisms:
- The status information "do_signal is handling a system call that may need
restarting" is itself carried in some register that can be accessed via
ptrace. This is e.g. on Intel the "orig_eax" register; on Sparc the kernel
defines a magic extra bit in the flags register for this purpose.
This allows GDB to manage that state: reset it when doing an inferior call,
and restore it after the call is finished.
- On s390, do_signal transparently handles this problem without requiring
GDB interaction, by performing system call restarting in the following
way: first, adjust the PC as necessary for restarting the call. Then,
call get_signal_to_deliver; and finally just continue execution at the
PC. This way, if GDB does not change the PC, everything is as before.
If GDB *does* change the PC, execution will simply continue there --
and once GDB restores the PC it saved at that point, it will automatically
point to the *restarted* system call. (There is the minor twist how to
handle system calls that do *not* need restarting -- do_signal will undo
the PC change in this case, after get_signal_to_deliver has returned, and
only if ptrace did not change the PC during that call.)
Because there does not appear to be any obvious register to carry the
syscall-restart information on ARM, we'd either have to introduce a new
artificial ptrace register just for that purpose, or else handle the issue
transparently like on s390. The patch below implements the second option;
using this patch makes the interrupt.exp test cases pass on ARM, with no
regression in the GDB test suite otherwise.
Cc: patches@linaro.org
Signed-off-by: Ulrich Weigand <ulrich.weigand@linaro.org>
Signed-off-by: Arnd Bergmann <arnd.bergmann@linaro.org>
Signed-off-by: Russell King <rmk+kernel@arm.linux.org.uk>
2011-05-03 18:32:55 +01:00
| | ( retval = = - ERESTARTSYS
2012-11-07 17:53:13 -05:00
& & ! ( ksig . ka . sa . sa_flags & SA_RESTART ) ) ) {
ARM: 6892/1: handle ptrace requests to change PC during interrupted system calls
GDB's interrupt.exp test cases currenly fail on ARM. The problem is how do_signal
handled restarting interrupted system calls:
The entry.S assembler code determines that we come from a system call; and that
information is passed as "syscall" parameter to do_signal. That routine then
calls get_signal_to_deliver [*] and if a signal is to be delivered, calls into
handle_signal. If a system call is to be restarted either after the signal
handler returns, or if no handler is to be called in the first place, the PC
is updated after the get_signal_to_deliver call, either in handle_signal (if
we have a handler) or at the end of do_signal (otherwise).
Now the problem is that during [*], the call to get_signal_to_deliver, a ptrace
intercept may happen. During this intercept, the debugger may change registers,
including the PC. This is done by GDB if it wants to execute an "inferior call",
i.e. the execution of some code in the debugged program triggered by GDB.
To this purpose, GDB will save all registers, allocate a stack frame, set up
PC and arguments as appropriate for the call, and point the link register to
a dummy breakpoint instruction. Once the process is restarted, it will execute
the call and then trap back to the debugger, at which point GDB will restore
all registers and continue original execution.
This generally works fine. However, now consider what happens when GDB attempts
to do exactly that while the process was interrupted during execution of a to-be-
restarted system call: do_signal is called with the syscall flag set; it calls
get_signal_to_deliver, at which point the debugger takes over and changes the PC
to point to a completely different place. Now get_signal_to_deliver returns
without a signal to deliver; but now do_signal decides it should be restarting
a system call, and decrements the PC by 2 or 4 -- so it now points to 2 or 4
bytes before the function GDB wants to call -- which leads to a subsequent crash.
To fix this problem, two things need to be supported:
- do_signal must be able to recognize that get_signal_to_deliver changed the PC
to a different location, and skip the restart-syscall sequence
- once the debugger has restored all registers at the end of the inferior call
sequence, do_signal must recognize that *now* it needs to restart the pending
system call, even though it was now entered from a breakpoint instead of an
actual svc instruction
This set of issues is solved on other platforms, usually by one of two
mechanisms:
- The status information "do_signal is handling a system call that may need
restarting" is itself carried in some register that can be accessed via
ptrace. This is e.g. on Intel the "orig_eax" register; on Sparc the kernel
defines a magic extra bit in the flags register for this purpose.
This allows GDB to manage that state: reset it when doing an inferior call,
and restore it after the call is finished.
- On s390, do_signal transparently handles this problem without requiring
GDB interaction, by performing system call restarting in the following
way: first, adjust the PC as necessary for restarting the call. Then,
call get_signal_to_deliver; and finally just continue execution at the
PC. This way, if GDB does not change the PC, everything is as before.
If GDB *does* change the PC, execution will simply continue there --
and once GDB restores the PC it saved at that point, it will automatically
point to the *restarted* system call. (There is the minor twist how to
handle system calls that do *not* need restarting -- do_signal will undo
the PC change in this case, after get_signal_to_deliver has returned, and
only if ptrace did not change the PC during that call.)
Because there does not appear to be any obvious register to carry the
syscall-restart information on ARM, we'd either have to introduce a new
artificial ptrace register just for that purpose, or else handle the issue
transparently like on s390. The patch below implements the second option;
using this patch makes the interrupt.exp test cases pass on ARM, with no
regression in the GDB test suite otherwise.
Cc: patches@linaro.org
Signed-off-by: Ulrich Weigand <ulrich.weigand@linaro.org>
Signed-off-by: Arnd Bergmann <arnd.bergmann@linaro.org>
Signed-off-by: Russell King <rmk+kernel@arm.linux.org.uk>
2011-05-03 18:32:55 +01:00
regs - > ARM_r0 = - EINTR ;
regs - > ARM_pc = continue_addr ;
}
}
2012-11-07 17:53:13 -05:00
handle_signal ( & ksig , regs ) ;
} else {
/* no handler */
restore_saved_sigmask ( ) ;
if ( unlikely ( restart ) & & regs - > ARM_pc = = restart_addr ) {
regs - > ARM_pc = continue_addr ;
return restart ;
}
2005-04-16 15:20:36 -07:00
}
2012-11-07 17:53:13 -05:00
return 0 ;
2005-04-16 15:20:36 -07:00
}
2012-07-19 17:48:21 +01:00
asmlinkage int
2012-07-19 17:47:55 +01:00
do_work_pending ( struct pt_regs * regs , unsigned int thread_flags , int syscall )
2005-04-16 15:20:36 -07:00
{
2015-08-20 16:13:37 +01:00
/*
* The assembly code enters us with IRQs off , but it hasn ' t
* informed the tracing code of that for efficiency reasons .
* Update the trace code with the current status .
*/
trace_hardirqs_off ( ) ;
2012-07-19 17:47:55 +01:00
do {
if ( likely ( thread_flags & _TIF_NEED_RESCHED ) ) {
schedule ( ) ;
} else {
if ( unlikely ( ! user_mode ( regs ) ) )
2012-07-19 17:48:21 +01:00
return 0 ;
2012-07-19 17:47:55 +01:00
local_irq_enable ( ) ;
2020-10-09 16:00:49 -06:00
if ( thread_flags & ( _TIF_SIGPENDING | _TIF_NOTIFY_SIGNAL ) ) {
2012-07-19 17:48:50 +01:00
int restart = do_signal ( regs , syscall ) ;
if ( unlikely ( restart ) ) {
2012-07-19 17:48:21 +01:00
/*
* Restart without handlers .
* Deal with it without leaving
* the kernel space .
*/
2012-07-19 17:48:50 +01:00
return restart ;
2012-07-19 17:48:21 +01:00
}
2012-07-19 17:47:55 +01:00
syscall = 0 ;
2014-03-07 11:23:04 -05:00
} else if ( thread_flags & _TIF_UPROBE ) {
uprobe_notify_resume ( regs ) ;
2012-07-19 17:47:55 +01:00
} else {
2022-02-09 12:20:45 -06:00
resume_user_mode_work ( regs ) ;
2012-07-19 17:47:55 +01:00
}
}
local_irq_disable ( ) ;
2021-11-29 13:06:47 +00:00
thread_flags = read_thread_flags ( ) ;
2012-07-19 17:47:55 +01:00
} while ( thread_flags & _TIF_WORK_MASK ) ;
2012-07-19 17:48:21 +01:00
return 0 ;
2005-04-16 15:20:36 -07:00
}
2013-07-24 00:29:18 +01:00
struct page * get_signal_page ( void )
{
2013-08-03 10:30:05 +01:00
unsigned long ptr ;
unsigned offset ;
struct page * page ;
void * addr ;
2013-07-24 00:29:18 +01:00
2013-08-03 10:30:05 +01:00
page = alloc_pages ( GFP_KERNEL , 0 ) ;
2013-07-24 00:29:18 +01:00
2013-08-03 10:30:05 +01:00
if ( ! page )
return NULL ;
2013-07-24 00:29:18 +01:00
2013-08-03 10:30:05 +01:00
addr = page_address ( page ) ;
2013-07-24 00:29:18 +01:00
2021-01-29 10:19:07 +00:00
/* Poison the entire page */
memset32 ( addr , __opcode_to_mem_arm ( 0xe7fddef1 ) ,
PAGE_SIZE / sizeof ( u32 ) ) ;
2013-08-03 10:30:05 +01:00
/* Give the signal return code some randomness */
2022-10-05 17:23:53 +02:00
offset = 0x200 + ( get_random_u16 ( ) & 0x7fc ) ;
2013-08-03 10:30:05 +01:00
signal_return_offset = offset ;
2013-07-24 00:29:18 +01:00
2021-01-29 10:19:07 +00:00
/* Copy signal return handlers into the page */
2013-08-03 10:30:05 +01:00
memcpy ( addr + offset , sigreturn_codes , sizeof ( sigreturn_codes ) ) ;
2013-07-24 00:29:18 +01:00
2021-01-29 10:19:07 +00:00
/* Flush out all instructions in this page */
ptr = ( unsigned long ) addr ;
flush_icache_range ( ptr , ptr + PAGE_SIZE ) ;
2013-07-24 00:29:18 +01:00
2013-08-03 10:30:05 +01:00
return page ;
2013-07-24 00:29:18 +01:00
}
2017-09-07 08:30:46 -07:00
2018-06-02 08:43:56 -04:00
# ifdef CONFIG_DEBUG_RSEQ
asmlinkage void do_rseq_syscall ( struct pt_regs * regs )
{
rseq_syscall ( regs ) ;
}
# endif
2021-04-29 21:07:33 +02:00
/*
* Compile - time assertions for siginfo_t offsets . Check NSIG * as well , as
* changes likely come with new fields that should be added below .
*/
static_assert ( NSIGILL = = 11 ) ;
static_assert ( NSIGFPE = = 15 ) ;
static_assert ( NSIGSEGV = = 9 ) ;
static_assert ( NSIGBUS = = 5 ) ;
static_assert ( NSIGTRAP = = 6 ) ;
static_assert ( NSIGCHLD = = 6 ) ;
static_assert ( NSIGSYS = = 2 ) ;
2021-05-04 11:25:22 -05:00
static_assert ( sizeof ( siginfo_t ) = = 128 ) ;
static_assert ( __alignof__ ( siginfo_t ) = = 4 ) ;
2021-04-29 21:07:33 +02:00
static_assert ( offsetof ( siginfo_t , si_signo ) = = 0x00 ) ;
static_assert ( offsetof ( siginfo_t , si_errno ) = = 0x04 ) ;
static_assert ( offsetof ( siginfo_t , si_code ) = = 0x08 ) ;
static_assert ( offsetof ( siginfo_t , si_pid ) = = 0x0c ) ;
static_assert ( offsetof ( siginfo_t , si_uid ) = = 0x10 ) ;
static_assert ( offsetof ( siginfo_t , si_tid ) = = 0x0c ) ;
static_assert ( offsetof ( siginfo_t , si_overrun ) = = 0x10 ) ;
static_assert ( offsetof ( siginfo_t , si_status ) = = 0x14 ) ;
static_assert ( offsetof ( siginfo_t , si_utime ) = = 0x18 ) ;
static_assert ( offsetof ( siginfo_t , si_stime ) = = 0x1c ) ;
static_assert ( offsetof ( siginfo_t , si_value ) = = 0x14 ) ;
static_assert ( offsetof ( siginfo_t , si_int ) = = 0x14 ) ;
static_assert ( offsetof ( siginfo_t , si_ptr ) = = 0x14 ) ;
static_assert ( offsetof ( siginfo_t , si_addr ) = = 0x0c ) ;
static_assert ( offsetof ( siginfo_t , si_addr_lsb ) = = 0x10 ) ;
static_assert ( offsetof ( siginfo_t , si_lower ) = = 0x14 ) ;
static_assert ( offsetof ( siginfo_t , si_upper ) = = 0x18 ) ;
static_assert ( offsetof ( siginfo_t , si_pkey ) = = 0x14 ) ;
static_assert ( offsetof ( siginfo_t , si_perf_data ) = = 0x10 ) ;
static_assert ( offsetof ( siginfo_t , si_perf_type ) = = 0x14 ) ;
signal: Deliver SIGTRAP on perf event asynchronously if blocked
With SIGTRAP on perf events, we have encountered termination of
processes due to user space attempting to block delivery of SIGTRAP.
Consider this case:
<set up SIGTRAP on a perf event>
...
sigset_t s;
sigemptyset(&s);
sigaddset(&s, SIGTRAP | <and others>);
sigprocmask(SIG_BLOCK, &s, ...);
...
<perf event triggers>
When the perf event triggers, while SIGTRAP is blocked, force_sig_perf()
will force the signal, but revert back to the default handler, thus
terminating the task.
This makes sense for error conditions, but not so much for explicitly
requested monitoring. However, the expectation is still that signals
generated by perf events are synchronous, which will no longer be the
case if the signal is blocked and delivered later.
To give user space the ability to clearly distinguish synchronous from
asynchronous signals, introduce siginfo_t::si_perf_flags and
TRAP_PERF_FLAG_ASYNC (opted for flags in case more binary information is
required in future).
The resolution to the problem is then to (a) no longer force the signal
(avoiding the terminations), but (b) tell user space via si_perf_flags
if the signal was synchronous or not, so that such signals can be
handled differently (e.g. let user space decide to ignore or consider
the data imprecise).
The alternative of making the kernel ignore SIGTRAP on perf events if
the signal is blocked may work for some usecases, but likely causes
issues in others that then have to revert back to interception of
sigprocmask() (which we want to avoid). [ A concrete example: when using
breakpoint perf events to track data-flow, in a region of code where
signals are blocked, data-flow can no longer be tracked accurately.
When a relevant asynchronous signal is received after unblocking the
signal, the data-flow tracking logic needs to know its state is
imprecise. ]
Fixes: 97ba62b27867 ("perf: Add support for SIGTRAP on perf events")
Reported-by: Dmitry Vyukov <dvyukov@google.com>
Signed-off-by: Marco Elver <elver@google.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Acked-by: Geert Uytterhoeven <geert@linux-m68k.org>
Tested-by: Dmitry Vyukov <dvyukov@google.com>
Link: https://lore.kernel.org/r/20220404111204.935357-1-elver@google.com
2022-04-04 13:12:04 +02:00
static_assert ( offsetof ( siginfo_t , si_perf_flags ) = = 0x18 ) ;
2021-04-29 21:07:33 +02:00
static_assert ( offsetof ( siginfo_t , si_band ) = = 0x0c ) ;
static_assert ( offsetof ( siginfo_t , si_fd ) = = 0x10 ) ;
static_assert ( offsetof ( siginfo_t , si_call_addr ) = = 0x0c ) ;
static_assert ( offsetof ( siginfo_t , si_syscall ) = = 0x10 ) ;
static_assert ( offsetof ( siginfo_t , si_arch ) = = 0x14 ) ;