mirror_ubuntu-kernels

mirror of https://git.proxmox.com/git/mirror_ubuntu-kernels.git synced 2026-01-04 16:21:11 +00:00

History

Waiman Long 0dc8c730c9 mutex: Make more scalable by doing less atomic operations In the __mutex_lock_common() function, an initial entry into the lock slow path will cause two atomic_xchg instructions to be issued. Together with the atomic decrement in the fast path, a total of three atomic read-modify-write instructions will be issued in rapid succession. This can cause a lot of cache bouncing when many tasks are trying to acquire the mutex at the same time. This patch will reduce the number of atomic_xchg instructions used by checking the counter value first before issuing the instruction. The atomic_read() function is just a simple memory read. The atomic_xchg() function, on the other hand, can be up to 2 order of magnitude or even more in cost when compared with atomic_read(). By using atomic_read() to check the value first before calling atomic_xchg(), we can avoid a lot of unnecessary cache coherency traffic. The only downside with this change is that a task on the slow path will have a tiny bit less chance of getting the mutex when competing with another task in the fast path. The same is true for the atomic_cmpxchg() function in the mutex-spin-on-owner loop. So an atomic_read() is also performed before calling atomic_cmpxchg(). The mutex locking and unlocking code for the x86 architecture can allow any negative number to be used in the mutex count to indicate that some tasks are waiting for the mutex. I am not so sure if that is the case for the other architectures. So the default is to avoid atomic_xchg() if the count has already been set to -1. For x86, the check is modified to include all negative numbers to cover a larger case. The following table shows the jobs per minutes (JPM) scalability data on an 8-node 80-core Westmere box with a 3.7.10 kernel. The numactl command is used to restrict the running of the high_systime workloads to 1/2/4/8 nodes with hyperthreading on and off. +-----------------+-----------+------------+----------+ \| Configuration \| Mean JPM \| Mean JPM \| % Change \| \| \| w/o patch \| with patch \| \| +-----------------+-----------------------------------+ \| \| User Range 1100 - 2000 \| +-----------------+-----------------------------------+ \| 8 nodes, HT on \| 36980 \| 148590 \| +301.8% \| \| 8 nodes, HT off \| 42799 \| 145011 \| +238.8% \| \| 4 nodes, HT on \| 61318 \| 118445 \| +51.1% \| \| 4 nodes, HT off \| 158481 \| 158592 \| +0.1% \| \| 2 nodes, HT on \| 180602 \| 173967 \| -3.7% \| \| 2 nodes, HT off \| 198409 \| 198073 \| -0.2% \| \| 1 node , HT on \| 149042 \| 147671 \| -0.9% \| \| 1 node , HT off \| 126036 \| 126533 \| +0.4% \| +-----------------+-----------------------------------+ \| \| User Range 200 - 1000 \| +-----------------+-----------------------------------+ \| 8 nodes, HT on \| 41525 \| 122349 \| +194.6% \| \| 8 nodes, HT off \| 49866 \| 124032 \| +148.7% \| \| 4 nodes, HT on \| 66409 \| 106984 \| +61.1% \| \| 4 nodes, HT off \| 119880 \| 130508 \| +8.9% \| \| 2 nodes, HT on \| 138003 \| 133948 \| -2.9% \| \| 2 nodes, HT off \| 132792 \| 131997 \| -0.6% \| \| 1 node , HT on \| 116593 \| 115859 \| -0.6% \| \| 1 node , HT off \| 104499 \| 104597 \| +0.1% \| +-----------------+------------+-----------+----------+ At low user range 10-100, the JPM differences were within +/-1%. So they are not that interesting. AIM7 benchmark run has a pretty large run-to-run variance due to random nature of the subtests executed. So a difference of less than +-5% may not be really significant. This patch improves high_systime workload performance at 4 nodes and up by maintaining transaction rates without significant drop-off at high node count. The patch has practically no impact on 1 and 2 nodes system. The table below shows the percentage time (as reported by perf record -a -s -g) spent on the __mutex_lock_slowpath() function by the high_systime workload at 1500 users for 2/4/8-node configurations with hyperthreading off. +---------------+-----------------+------------------+---------+ \| Configuration \| %Time w/o patch \| %Time with patch \| %Change \| +---------------+-----------------+------------------+---------+ \| 8 nodes \| 65.34% \| 0.69% \| -99% \| \| 4 nodes \| 8.70% \| 1.02% \| -88% \| \| 2 nodes \| 0.41% \| 0.32% \| -22% \| +---------------+-----------------+------------------+---------+ It is obvious that the dramatic performance improvement at 8 nodes was due to the drastic cut in the time spent within the __mutex_lock_slowpath() function. The table below show the improvements in other AIM7 workloads (at 8 nodes, hyperthreading off). +--------------+---------------+----------------+-----------------+ \| Workload \| mean % change \| mean % change \| mean % change \| \| \| 10-100 users \| 200-1000 users \| 1100-2000 users \| +--------------+---------------+----------------+-----------------+ \| alltests \| +0.6% \| +104.2% \| +185.9% \| \| five_sec \| +1.9% \| +0.9% \| +0.9% \| \| fserver \| +1.4% \| -7.7% \| +5.1% \| \| new_fserver \| -0.5% \| +3.2% \| +3.1% \| \| shared \| +13.1% \| +146.1% \| +181.5% \| \| short \| +7.4% \| +5.0% \| +4.2% \| +--------------+---------------+----------------+-----------------+ Signed-off-by: Waiman Long <Waiman.Long@hp.com> Reviewed-by: Davidlohr Bueso <davidlohr.bueso@hp.com> Reviewed-by: Rik van Riel <riel@redhat.com> Cc: Linus Torvalds <torvalds@linux-foundation.org> Cc: Andrew Morton <akpm@linux-foundation.org> Cc: Peter Zijlstra <a.p.zijlstra@chello.nl> Cc: Thomas Gleixner <tglx@linutronix.de> Cc: Chandramouleeswaran Aswin <aswin@hp.com> Cc: Norton: Scott J <scott.norton@hp.com> Cc: Paul E. McKenney <paulmck@linux.vnet.ibm.com> Cc: David Howells <dhowells@redhat.com> Cc: Dave Jones <davej@redhat.com> Cc: Clark Williams <williams@redhat.com> Cc: Peter Zijlstra <peterz@infradead.org> Link: http://lkml.kernel.org/r/1366226594-5506-3-git-send-email-Waiman.Long@hp.com Signed-off-by: Ingo Molnar <mingo@kernel.org>		2013-04-19 09:33:35 +02:00
..
boot	x86: Fix rebuild with EFI_STUB enabled	2013-04-05 13:59:23 -07:00
configs	x86: Default to ARCH=x86 to avoid overriding CONFIG_64BIT	2012-12-20 14:37:18 -08:00
crypto	Merge git://git.kernel.org/pub/scm/linux/kernel/git/herbert/crypto-2.6	2013-02-25 15:56:15 -08:00
ia32	Merge branch 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/signal	2013-03-02 08:34:06 -08:00
include	mutex: Make more scalable by doing less atomic operations	2013-04-19 09:33:35 +02:00
kernel	Merge branch 'x86/urgent' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2013-03-24 10:10:34 -07:00
kvm	KVM: Allow cross page reads and writes from cached translations.	2013-04-07 13:05:35 +03:00
lguest	Merge branch 'x86-mm-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2013-02-21 18:06:55 -08:00
lib	x86-64: Fix the failure case in copy_user_handle_tail()	2013-03-18 11:32:03 -07:00
math-emu
mm	x86: Do not try to sync identity map for non-mapped pages	2013-03-07 13:23:28 -08:00
net	x86: bpf_jit_comp: add pkt_type support	2013-01-30 22:38:34 -05:00
oprofile	oprofile, x86: Fix wrapping bug in op_x86_get_ctrl()	2012-10-15 14:38:24 +02:00
pci	Bug-fixes:	2013-03-03 14:22:53 -08:00
platform	Merge branch 'x86-efi-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2013-02-27 16:17:42 -08:00
power	perf,x86: fix kernel crash with PEBS/BTS after suspend/resume	2013-03-15 09:26:35 -07:00
realmode	Merge remote-tracking branch 'origin/x86/mm' into x86/mm2	2013-02-01 02:28:36 -08:00
syscalls	fix compat truncate/ftruncate	2013-02-25 09:24:55 -05:00
tools	Merge branch 'perf-urgent-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2013-02-05 07:57:09 +11:00
um	Merge branch 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/signal	2013-02-23 18:50:11 -08:00
vdso	timers/x86/hpet: Use HPET_COUNTER to specify the hpet counter in vread_hpet()	2013-02-15 12:13:18 +01:00
video	x86: Use vga_default_device() when determining whether an fb is primary	2012-04-24 09:50:17 +01:00
xen	xen/mmu: Move the setting of pvops.write_cr3 to later phase in bootup.	2013-03-27 12:06:03 -04:00
.gitignore
Kbuild	x86, realmode: realmode.bin infrastructure	2012-05-08 11:41:48 -07:00
Kconfig	Select VIRT_TO_BUS directly where needed	2013-03-12 11:16:40 -07:00
Kconfig.cpu	x86, 386 removal: Document Nx586 as a 386 and thus unsupported	2012-11-29 13:28:39 -08:00
Kconfig.debug	x86/tlb: add tlb_flushall_shift knob into debugfs	2012-06-27 19:29:10 -07:00
Makefile	Merge branch 'x86-build-for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip	2013-02-19 19:12:03 -08:00
Makefile_32.cpu	x86, 386 removal: Remove CONFIG_M386 from Kconfig	2012-11-29 13:23:01 -08:00
Makefile.um	um: fix linker script generation	2012-04-09 13:59:00 -04:00