android_kernel_xiaomi_sdm845

History

Jeremy Fitzhardinge b4ecc12699 x86: Fix performance regression caused by paravirt_ops on native kernels Xiaohui Xin and some other folks at Intel have been looking into what's behind the performance hit of paravirt_ops when running native. It appears that the hit is entirely due to the paravirtualized spinlocks introduced by: \| commit `8efcbab674` \| Date: Mon Jul 7 12:07:51 2008 -0700 \| \| paravirt: introduce a "lock-byte" spinlock implementation The extra call/return in the spinlock path is somehow causing an increase in the cycles/instruction of somewhere around 2-7% (seems to vary quite a lot from test to test). The working theory is that the CPU's pipeline is getting upset about the call->call->locked-op->return->return, and seems to be failing to speculate (though I haven't seen anything definitive about the precise reasons). This doesn't entirely make sense, because the performance hit is also visible on unlock and other operations which don't involve locked instructions. But spinlock operations clearly swamp all the other pvops operations, even though I can't imagine that they're nearly as common (there's only a .05% increase in instructions executed). If I disable just the pv-spinlock calls, my tests show that pvops is identical to non-pvops performance on native (my measurements show that it is actually about .1% faster, but Xiaohui shows a .05% slowdown). Summary of results, averaging 10 runs of the "mmperf" test, using a no-pvops build as baseline: nopv Pv-nospin Pv-spin CPU cycles 100.00% 99.89% 102.18% instructions 100.00% 100.10% 100.15% CPI 100.00% 99.79% 102.03% cache ref 100.00% 100.84% 100.28% cache miss 100.00% 90.47% 88.56% cache miss rate 100.00% 89.72% 88.31% branches 100.00% 99.93% 100.04% branch miss 100.00% 103.66% 107.72% branch miss rt 100.00% 103.73% 107.67% wallclock 100.00% 99.90% 102.20% The clear effect here is that the 2% increase in CPI is directly reflected in the final wallclock time. (The other interesting effect is that the more ops are out of line calls via pvops, the lower the cache access and miss rates. Not too surprising, but it suggests that the non-pvops kernel is over-inlined. On the flipside, the branch misses go up correspondingly...) So, what's the fix? Paravirt patching turns all the pvops calls into direct calls, so _spin_lock etc do end up having direct calls. For example, the compiler generated code for paravirtualized _spin_lock is: <_spin_lock+0>: mov %gs:0xb4c8,%rax <_spin_lock+9>: incl 0xffffffffffffe044(%rax) <_spin_lock+15>: callq 0xffffffff805a5b30 <_spin_lock+22>: retq The indirect call will get patched to: <_spin_lock+0>: mov %gs:0xb4c8,%rax <_spin_lock+9>: incl 0xffffffffffffe044(%rax) <_spin_lock+15>: callq <__ticket_spin_lock> <_spin_lock+20>: nop; nop / or whatever 2-byte nop */ <_spin_lock+22>: retq One possibility is to inline _spin_lock, etc, when building an optimised kernel (ie, when there's no spinlock/preempt instrumentation/debugging enabled). That will remove the outer call/return pair, returning the instruction stream to a single call/return, which will presumably execute the same as the non-pvops case. The downsides arel 1) it will replicate the preempt_disable/enable code at eack lock/unlock callsite; this code is fairly small, but not nothing; and 2) the spinlock definitions are already a very heavily tangled mass of #ifdefs and other preprocessor magic, and making any changes will be non-trivial. The other obvious answer is to disable pv-spinlocks. Making them a separate config option is fairly easy, and it would be trivial to enable them only when Xen is enabled (as the only non-default user). But it doesn't really address the common case of a distro build which is going to have Xen support enabled, and leaves the open question of whether the native performance cost of pv-spinlocks is worth the performance improvement on a loaded Xen system (10% saving of overall system CPU when guests block rather than spin). Still it is a reasonable short-term workaround. [ Impact: fix pvops performance regression when running native ] Analysed-by: "Xin Xiaohui" <xiaohui.xin@intel.com> Analysed-by: "Li Xin" <xin.li@intel.com> Analysed-by: "Nakajima Jun" <jun.nakajima@intel.com> Signed-off-by: Jeremy Fitzhardinge <jeremy.fitzhardinge@citrix.com> Acked-by: H. Peter Anvin <hpa@zytor.com> Cc: Nick Piggin <npiggin@suse.de> Cc: Xen-devel <xen-devel@lists.xensource.com> LKML-Reference: <4A0B62F7.5030802@goop.org> [ fixed the help text ] Signed-off-by: Ingo Molnar <mingo@elte.hu>		2009-05-15 20:07:42 +02:00
..
alpha	alpha: binfmt_aout fix	2009-05-02 15:36:10 -07:00
arm	[ARM] 5507/1: support R_ARM_MOVW_ABS_NC and MOVT_ABS relocation types	2009-05-07 17:21:01 +01:00
avr32	avr32: drop unused CLEAN_FILES	2009-05-01 10:54:00 +02:00
blackfin	clocksource: pass clocksource to read() callback	2009-04-21 13:41:47 -07:00
cris	tty: Use the generic RS485 ioctl on CRIS	2009-04-07 08:44:05 -07:00
frv	FRV: Use __INIT macro instead of .text.init.	2009-04-27 19:46:30 -07:00
h8300	Get rid of final remnants of include/asm-$(ARCH)	2009-04-17 09:59:27 -07:00
ia64	[IA64] xen_domu_defconfig: fix build issues/warnings	2009-05-05 11:43:13 -07:00
m32r	m32r: use __stringify() macro in assembler.h	2009-05-02 22:38:21 +09:00
m68k	m68k: arch/m68k/kernel/sun3-head.S needs <linux/init.h>	2009-04-28 16:07:18 -07:00
m68knommu	Merge branch 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/gerg/m68knommu	2009-04-24 08:45:53 -07:00
microblaze	Merge branch 'fixes-for-linus' of git://git.monstr.eu/linux-2.6-microblaze	2009-05-08 16:24:25 -07:00
mips	clocksource: pass clocksource to read() callback	2009-04-21 13:41:47 -07:00
mn10300	mn10300: convert to use __HEAD and HEAD_TEXT macros.	2009-04-26 09:20:38 -07:00
parisc	Merge git://git.kernel.org/pub/scm/linux/kernel/git/kyle/parisc-2.6	2009-04-03 09:52:04 -07:00
powerpc	Merge branch 'merge' of git://git.kernel.org/pub/scm/linux/kernel/git/benh/powerpc	2009-05-05 08:25:37 -07:00
s390	s390: convert to use __HEAD and HEAD_TEXT macros.	2009-04-26 09:20:39 -07:00
sh	sh: Use __INIT macro instead of .text.init.	2009-04-27 19:51:58 -07:00
sparc	sparc: cleanup references to deprecated .text.init* sections.	2009-04-27 19:51:58 -07:00
um	uml: kill a kconfig warning	2009-04-21 13:41:50 -07:00
x86	x86: Fix performance regression caused by paravirt_ops on native kernels	2009-05-15 20:07:42 +02:00
xtensa	xtensa: convert to use __HEAD and HEAD_TEXT macros.	2009-04-26 09:20:38 -07:00
.gitignore
Kconfig	mutex: have non-spinning mutexes on s390 by default	2009-04-09 19:28:24 +02:00