mesa.git - Unnamed repository; edit this file 'description' to name the repository.

	Commit message (Collapse)	Author	Age	Files	Lines
*	gallium/radeon: Correctly translate colorswaps for big endian	Oded Gabbay	2016-02-23	1	-0/+11
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The current code in r600_translate_colorswap uses the swizzle information to determine which colorswap to use. This works for BE & LE when the nr_channels is <4, but when nr_channels==4 (e.g. PIPE_FORMAT_A8R8G8B8_UNORM), this method can not be used for both BE and LE, because the swizzle info is the same for both of them. As a result, r600g doesn't support 24bit color formats, only 16bit, which forces the user to choose 16bit color in X server. This patch fixes this bug by separating the checks for LE and BE and adapting the swizzle conditions in the BE part of the checks. Tested on an Evergreen GPU (Cedar GL FirePro 2270) running inside POWER7 Big-Endian Machine. Signed-off-by: Oded Gabbay <[email protected]> CC: "11.2" "11.1" <[email protected]> Reviewed-by: Marek Olšák <[email protected]>
*	tgsi/scan: handle holes between VS inputs, assert-fail in other cases	Marek Olšák	2016-02-23	1	-1/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	"st/mesa: overhaul vertex setup for clearing, glDrawPixels, glBitmap" added a vertex shader declaring IN[0] and IN[2], but not IN[1]. Drivers relying on tgsi_shader_info can't handle holes in declarations, because tgsi_shader_info doesn't track that. This is just a quick workaround meant for stable that will work for vertex shaders. This fixes radeonsi DrawPixels and CopyPixels crashes. Cc: [email protected] Reviewed-by: Brian Paul <[email protected]>
*	nvc0: rename 3d binding points to NVC0_BIND_3D_XXX	Samuel Pitoiset	2016-02-22	9	-63/+64
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: rename 3d dirty flags to NVC0_NEW_3D_XXX	Samuel Pitoiset	2016-02-22	8	-133/+133
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: prefix compute macros with _CP_ instead of _COMPUTE_	Samuel Pitoiset	2016-02-22	4	-4/+4
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: rename NVXX_COMPUTE to NVXX_CP	Samuel Pitoiset	2016-02-22	5	-117/+117
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: rename nvc0_context::dirty to nvc0_context::dirty_3d	Samuel Pitoiset	2016-02-22	8	-64/+64
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0/ir: add missing emission of locked load predicate	Samuel Pitoiset	2016-02-22	1	-0/+7
\| \| \| \| \| \| \| \| \|	Like unlocked store on shared memory, locked store can fail and the second dest which is a predicate must be emitted. Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]> Cc: [email protected]
*	nvc0/ir: add ld lock/st unlock emission on GK104	Samuel Pitoiset	2016-02-22	1	-10/+25
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nv50/ir: restore OP_SELP to be a regular instruction	Samuel Pitoiset	2016-02-22	4	-14/+14
\| \| \| \| \| \| \| \| \| \| \|	Actually OP_SELP doesn't need to be a compare instruction. Instead we just need to set the NOT modifier when building the instruction. While we are at it, fix the dst register type and use a GPR. Suggested by Ilia Mirkin. Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	svga: unbind index buffer when drawing non-indexed primitives	Brian Paul	2016-02-22	1	-0/+10
\| \| \| \| \| \| \| \| \|	Silences a warning reported by the svga3d device. v2: also null-out the index buffer pointer Reviewed-by: Sinclair Yeh <[email protected]> Reviewed-by: Charmaine Lee <[email protected]>
*	nouveau: update the Makefile.sources list11.2-branchpoint	Emil Velikov	2016-02-22	1	-2/+3
\| \| \| \| \| \|	Reflect the nv50->g80 change and the new gm107_texture header. Signed-off-by: Emil Velikov <[email protected]>
*	radeonsi: implement binary shaders & shader cache in memory (v2)	Marek Olšák	2016-02-21	5	-7/+259
\| \| \| \| \| \| \|	v2: handle _mesa_hash_table_insert failure other cosmetic changes Reviewed-by: Nicolai Hähnle <[email protected]>
*	gallium/radeon: remove unused radeon_shader_binary_free_* functions	Marek Olšák	2016-02-21	2	-33/+0
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: make radeon_shader_reloc name string fixed-sized	Marek Olšák	2016-02-21	2	-6/+3
\| \| \| \| \| \|	This will simplify implementations of binary shaders. Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: move some struct si_shader members to new struct si_shader_info	Marek Olšák	2016-02-21	3	-68/+71
\| \| \| \| \| \|	This will be part of shader binaries. Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: use smaller types for some si_shader members	Marek Olšák	2016-02-21	2	-3/+8
\| \| \| \| \| \| \| \|	in order to decrease the shader size for a shader cache. v2: add & use SI_MAX_VS_OUTPUTS Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: enable compiling one variant per shader	Marek Olšák	2016-02-21	3	-1/+5
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Shader stats from VERDE: Default scheduler: Totals: SGPRS: 491272 -> 488672 (-0.53 %) VGPRS: 289980 -> 311093 (7.28 %) Code Size: 11091656 -> 11219948 (1.16 %) bytes LDS: 97 -> 97 (0.00 %) blocks Scratch: 1732608 -> 2246656 (29.67 %) bytes per wave Max Waves: 78063 -> 77352 (-0.91 %) Wait states: 0 -> 0 (0.00 %) Looking at some of the worst regressions, I get: - The VGPR increase seems to be caused by the fact that if PS has used less than 16 VGPRs, now it will always use 16 VGPRs and sometimes even 20. However, the wave count remains at 10 if VGPRs <= 24, so no harm there. - The scratch increase seems to be caused by SGPR spilling. The unnecessary SGPR spilling has been an ongoing issue with the compiler and it's completely fixable by rematerializing s_loads or reordering instructions. SI scheduler: Totals: SGPRS: 374848 -> 374576 (-0.07 %) VGPRS: 284456 -> 307515 (8.11 %) Code Size: 11433068 -> 11535452 (0.90 %) bytes LDS: 97 -> 97 (0.00 %) blocks Scratch: 509952 -> 522240 (2.41 %) bytes per wave Max Waves: 79456 -> 78217 (-1.56 %) Wait states: 0 -> 0 (0.00 %) VGPRs - same story as before. The SI scheduler doesn't spill SGPRs so much and generally spills way less than the default scheduler. (522240 spills vs 2246656 spills) Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: print full shader name before disassembly	Marek Olšák	2016-02-21	1	-1/+33
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: compile non-GS middle parts of shaders immediately if enabled	Marek Olšák	2016-02-21	3	-18/+87
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Still disabled. Only prologs & epilogs are compiled in draw calls, but each variant of those is compiled only once per process. VS is always compiled as hw VS. TES is always compiled as hw VS. LS and ES stages are always compiled on demand. Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: rework polygon stippling for PS prolog	Marek Olšák	2016-02-21	1	-39/+110
\| \| \| \| \| \|	Don't use the pstipple module. Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add PS prolog	Marek Olšák	2016-02-21	5	-2/+345
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add PS epilog	Marek Olšák	2016-02-21	4	-2/+297
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add TCS epilog	Marek Olšák	2016-02-21	4	-13/+155
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add VS epilog	Marek Olšák	2016-02-21	4	-11/+171
\| \| \| \| \| \| \| \| \|	It only exports the primitive ID. Also used by TES when it's compiled as VS. The VS input location of the primitive ID input is v2. Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add VS prolog	Marek Olšák	2016-02-21	4	-1/+267
\| \| \| \| \| \|	This is disabled with use_monolithic_shaders = true. Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: first bits for non-monolithic shaders	Marek Olšák	2016-02-21	4	-14/+45
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add code for dumping all shader parts together (v2)	Marek Olšák	2016-02-21	1	-12/+34
\| \| \| \| \| \|	v2: unify some code into si_get_shader_binary_size Reviewed-by: Michel Dänzer <[email protected]>
*	radeonsi: add code for combining and uploading shaders from 3 shader parts	Marek Olšák	2016-02-21	2	-8/+36
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: fail compilation if non-GS non-CS shaders have rodata	Marek Olšák	2016-02-21	1	-0/+13
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: separate 2 pieces of code from create_function	Marek Olšák	2016-02-21	1	-31/+51
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add samplemask parameter to si_export_mrt_color	Marek Olšák	2016-02-21	1	-3/+7
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: add start_instance parameter to get_instance_index_for_fetch	Marek Olšák	2016-02-21	1	-4/+6
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: separate out shader key bits for prologs & epilogs	Marek Olšák	2016-02-21	4	-100/+140
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: compute how many input VGPRs fragment shaders have	Marek Olšák	2016-02-21	2	-0/+43
\| \| \| \|	Reviewed-by: Nicolai Hähnle <[email protected]>
*	radeonsi: compute how many input SGPRs and VGPRs shaders have	Marek Olšák	2016-02-21	2	-0/+34
\| \| \| \| \| \| \|	Prologs (shader binaries inserted before the API shader binary) need to know this, so that they won't change the input registers unintentionally. Reviewed-by: Nicolai Hähnle <[email protected]>
*	gallium/radeon: add basic code for setting shader return values	Marek Olšák	2016-02-21	4	-8/+21
\| \| \| \| \| \|	LLVMBuildInsertValue will be used on return_value. Reviewed-by: Nicolai Hähnle <[email protected]>
*	nvc0: enable compute shaders on Fermi	Samuel Pitoiset	2016-02-21	1	-1/+3
\| \| \| \| \| \| \| \|	Kepler compute support is really different than Fermi and it's not ready yet. Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nv50/ir: add atomics support on shared memory for Fermi	Samuel Pitoiset	2016-02-21	2	-2/+102
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Changes from v3: - move the previous OP_SELP change to the previous commit Changes from v2: - make sure the op is OP_SELP when emitting the predicate and add one assert - use bld.getSSA() for mkOp2() - add cross edge between tryLockAndSetBB and joinBB Signed-off-by: Samuel Pitoiset <[email protected]> Acked-by: Ilia Mirkin <[email protected]>
*	nv50/ir: make OP_SELP a compare instruction	Samuel Pitoiset	2016-02-21	3	-10/+19
\| \| \| \| \| \| \| \| \| \|	This OP_SELP insn will be used to handle compare and swap subops. Changes from v2: - fix logic for GK110+ Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nv50/ir: add lock/unlock subops for load/store	Samuel Pitoiset	2016-02-21	3	-2/+26
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nv50/ir: use s[] addr space for shared buffers	Samuel Pitoiset	2016-02-21	1	-11/+30
\| \| \| \| \| \| \| \| \| \| \| \|	Shared memory address space (FILE_MEMORY_SHARED) must be used instead of global memory when a shared memory area is declared. Changes from v2: - oops, do not remove TGSI_FILE_BUFFER in a switch in nv50_ir_from_tgsi.cpp Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: reduce likelihood of collision for real buffers on Fermi	Samuel Pitoiset	2016-02-21	1	-2/+2
\| \| \| \| \| \| \| \| \| \| \|	Reduce likelihood of collision with real buffers by placing the hole at the top of the 4G area. This fixes some indirect draw+compute tests with large buffers. Suggested by Ilia Mirkin. Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: invalidate compute state when switching pipe contexts	Samuel Pitoiset	2016-02-21	1	-1/+2
\| \| \| \| \|	Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: add support for indirect compute on Fermi	Samuel Pitoiset	2016-02-21	6	-20/+81
\| \| \| \| \| \| \| \| \| \| \| \|	When indirect compute is used, the size of the grid (in blocks) is stored as three integers inside a buffer. This requires a macro to set up GRIDDIM_YX and GRIDDIM_Z. Changes from v2: - do not launch the grid if the number of groups for a dimension is 0 Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: bind textures/samplers for compute on Fermi	Samuel Pitoiset	2016-02-21	3	-8/+57
\| \| \| \| \| \| \| \| \| \| \|	Textures and samplers don't seem to be aliased between COMPUTE and 3D. Changes from v2: - refactor the code to share (almost) the same logic between 3d and compute Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: bind shader buffers for compute on Fermi	Samuel Pitoiset	2016-02-21	5	-7/+56
\| \| \| \| \| \| \| \|	This is loosely based on 3D. Shader buffers are bound on c15 (the driver constbuf) at offset 0x200. Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: bind driver constbuf for compute on Fermi	Samuel Pitoiset	2016-02-21	5	-0/+29
\| \| \| \| \| \| \| \| \| \| \|	Changes from v3: - add new validation state for COMPUTE driver constbuf Changes from v2: - always bind the driver consts even if user params come in via clover Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: add a new validation state for 3D driver constbuf	Samuel Pitoiset	2016-02-21	2	-0/+19
\| \| \| \| \| \| \| \| \|	This will be used to invalidate 3D driver constbuf when using COMPUTE and vice-versa. This is needed because this CB contains a bunch of useful information like the addrs of shader buffers. Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	nvc0: bind constant buffers for compute on Fermi	Samuel Pitoiset	2016-02-21	5	-13/+81
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Loosely based on 3D. Changs from v3: - invalidate COMPUTE CBs after validating 3D CBs because they are aliased Changes from v2: - get rid of the 's' param to nvc0_cb_bo_push() because it doesn't matter to upload constbufs for compute using the 3d chan Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>