mesa.git - Unnamed repository; edit this file 'description' to name the repository.

	Commit message (Collapse)	Author	Age	Files	Lines
*	nvc0: respect edgeflag attribute width	Ilia Mirkin	2015-10-23	1	-7/+33
\| \| \| \| \| \| \| \| \| \| \| \|	The edgeflag comes in as ubyte with glEdgeFlagPointer but as float with plain immediate glEdgeFlag. Avoid reading bytes that weren't meant for the edgeflag in the pointer case. Fixes intermittent failures with gl-2.0-edgeflag piglit (and valgrind complaints about reading uninitialized memory). Signed-off-by: Ilia Mirkin <[email protected]> Cc: [email protected]
*	vc4: Convert blending to being done in 4x8 unorm normally.	Eric Anholt	2015-10-23	5	-51/+276
\| \| \| \| \| \| \| \| \| \| \| \| \|	We can't do this all the time, because you want blending to be done in linear space, and sRGB would lose too much precision being done in 4x8. The win on instructions is pretty huge when you can, though. total uniforms in shared programs: 32065 -> 32168 (0.32%) uniforms in affected programs: 327 -> 430 (31.50%) total instructions in shared programs: 92644 -> 89830 (-3.04%) instructions in affected programs: 15580 -> 12766 (-18.06%) Improves openarena performance at 1920x1080 from 10.7fps to 11.2fps.
*	vc4: Add QIR/QPU support for the 8-bit vector instructions.	Eric Anholt	2015-10-23	4	-0/+45
\|
*	vc4: Don't try to CSE non-SSA instructions.	Eric Anholt	2015-10-23	1	-0/+1
\| \| \| \| \| \| \|	This can happen when we're doing destination packing -- we don't know what's in the rest of the register. Signed-off-by: Eric Anholt <[email protected]>
*	vc4: Add dumping of VC4_PACKET_GL_INDEXED_PRIMITIVE.	Eric Anholt	2015-10-23	1	-1/+22
\|
*	vc4: Add a workaround for HW-2116 (state counter wrap fails).	Eric Anholt	2015-10-23	3	-6/+40
\| \| \| \| \| \|	I haven't proven that this happens (I've got other GPU hangs in the way), but the closed driver also does this and it's documented as an errata.
*	vc4: Fix missing \n in a perf_debug().	Eric Anholt	2015-10-23	1	-1/+1
\|
*	vc4: Use Rob's NIR-based user clip lowering.	Eric Anholt	2015-10-23	4	-69/+14
\|
*	vc4: Also dump the decimation mode for resolved stores.	Eric Anholt	2015-10-23	1	-2/+4
\|
*	vc4: Use VC4_GET_FIELD and other defines in dumping VC4_RENDER_CONFIG.	Eric Anholt	2015-10-23	1	-10/+10
\|
*	vc4: Add a sentinel after simulator buffers for buffer overflow detection.	Eric Anholt	2015-10-23	1	-1/+11
\| \| \| \| \| \| \| \| \|	This is a little bit like the mprotect-based fencing I've experimented with, but it's simple and low overhead. The downside is that only catches writes, not reads. It didn't catch any bad writes on a current piglit run, but may be useful in the future.
*	ilo: add support for scratch spaces	Chia-I Wu	2015-10-23	10	-16/+133
\| \| \| \| \|	When a kernel reports a non-zero per-thread scratch space size, make sure the hardware state is correctly set up, and a scratch bo is allocated.
*	ilo: fix scratch space setup in core	Chia-I Wu	2015-10-23	11	-133/+327
\| \| \| \| \| \|	Move scratch_size out of ilo_state_shader_kernel_info and ilo_state_compute_interface_info. A scratch space is shared by all kernels/interfaces. Update builder to emit relocs for scratch bos.
*	virgl/vtest: add vtest driver	Dave Airlie	2015-10-23	1	-1/+2
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	virgl/vtest is a swrast driver that allows the virgl acceleration to be tested without having a virtual machine. The backend has a unix socket server that this connects to. This is run by setting LIBGL_ALWAYS_SOFTWARE=y GALLIUM_DRIVER=virpipe In this mode all renderering is sent over a socket to the remote renderer, and the results are readback and copies to the screen using drisw. This works well enough to develop new features and to help debug. Signed-off-by: Dave Airlie <[email protected]>
*	virgl: add driver for virtio-gpu 3D (v2)	Dave Airlie	2015-10-23	18	-0/+4492
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	virgl is the 3D acceleration backend for the virtio-gpu shipping with qemu. The 3D acceleration is designed around gallium and TGSI as the virtualisation layer. The backend renderer translates the virgl interface into OpenGL currently. This is the initial import of the driver to mesa. The kernel driver portions are lined up for drm-next. Currently this driver supports up to GL3.3 and some misc extensions if the host driver exposes it. It is planned to iterate the virgl API to new GL levels as mesa host drivers gain features. v2: fix resource tracking across flushes to avoid ->bind hack in mapping. consolidate mapping and waiting code for transfers. use u_range for dirt tracking. handle larger shaders in protocol. include virtgpu_drm.h in mesa for now. add translation layer for gallium tgsi to virgl tgsi. Signed-off-by: Dave Airlie <[email protected]>
*	svga: Condition preemptive flush on draw emission	Sinclair Yeh	2015-10-22	3	-0/+15
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	On ultra high resolution modes, the preemptive flush flag can be set midway through command submission, a condition that cannot be recovered from a flush-retry, causing rendering artifacts. This patch prevents a preemtive_flush until a draw has been emitted. Signed-off-by: Sinclair Yeh <[email protected]> Reviewed-by: Thomas Hellstrom <[email protected]> Reviewed-by: Charmaine Lee <[email protected]> Reviewed-by: Brian Paul <[email protected]>
*	svga: try to avoid index generation for some primitive types	Brian Paul	2015-10-22	1	-0/+14
\| \| \| \| \| \| \| \| \| \|	The svga device doesn't directly support quads, quad strips or polygons so we have to convert those types to indexed triangle lists. But we can sometimes avoid that if we're drawing flat/constant-colored prims and we don't have to worry about provoking vertex. Reviewed-by: Charmaine Lee <[email protected]> Reviewed-by: José Fonseca <[email protected]>
*	svga: avoid provoking vertex conversion when possible	Brian Paul	2015-10-22	1	-1/+14
\| \| \| \| \| \| \| \| \| \| \| \|	Provoking vertex comes into play when doing flat shading. But if we know that all fragments in a primitive are the same color, the provoking vertex doesn't matter. Check for that case and use whichever provoking vertex convention is supported by the device. This avoids generating an index buffer to do the PV conversion. Reviewed-by: Charmaine Lee <[email protected]> Reviewed-by: José Fonseca <[email protected]>
*	svga: detect constant color writes in fragment shaders	Brian Paul	2015-10-22	5	-2/+77
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Examine the fragment shader to try to detect TGSI shaders which use "MOV OUT[0], CONST[i]" to write a constant value for the fragment color. In this case, all fragments will have the same color (unless blending is enabled). This is a common case for OpenGL code such as: glColor(), glBegin(), glVertex(), ..., glEnd() when lighting/fog/etc are disabled. In this case, the Mesa/gallium state tracker actually generates a simple "MOV OUT[0], CONST[i]" fragment shader. This will be used by the next commit to avoid provoking vertex conversion (creating/rewriting an index buffer) when drawing flat-shaded primitives. Reviewed-by: Charmaine Lee <[email protected]> Reviewed-by: José Fonseca <[email protected]>
*	radeon/uvd: don't expose HEVC on old UVD hw (v3)	Alex Deucher	2015-10-22	1	-32/+18
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	The section for UVD 2 and older was not updated when HEVC support was added. Reported by Kano on irc. v2: integrate the UVD2 and older checks into the main switch statement. v3: handle encode checking as well. Encode is already checked in the top case statement, so drop encode checks in the lower case statement. Reviewed-by: Christian König <[email protected]> Signed-off-by: Alex Deucher <[email protected]> Cc: [email protected]
*	ilo: make sure there is HiZ before resolving	Chia-I Wu	2015-10-22	1	-2/+4
\| \| \| \|	We do not want to perform a depth resolve on an MCS enabled surface.
*	ilo: fix max thread count for HS on Gen8	Chia-I Wu	2015-10-22	1	-3/+5
\| \| \| \|	It is in DW2 on Gen8.
*	svga: fix clip plane regression after recent tgsi_scan change	Brian Paul	2015-10-21	1	-2/+2
\| \| \| \| \| \| \| \| \|	Before the change "tgsi/scan: use properties for clip/cull distance writemasks", the tgsi_shader_info::num_written_clipdistance field was a multiple of four, now it's an accurate count. In the svga driver, we need a minor change to the loop test. Reviewed-by: Charmaine Lee <[email protected]>
*	svga: add switch case for PIPE_SHADER_CAP_MAX_UNROLL_ITERATIONS_HINT	Brian Paul	2015-10-20	1	-0/+2
\| \| \| \| \| \| \| \|	A third instance of this was needed but missed in the previous commit. Return 32 as for the two other cases. Reviewed-by: Roland Scheidegger <[email protected]> Reviewed-by: Charmaine Lee <[email protected]>
*	gallium: add PIPE_SHADER_CAP_MAX_UNROLL_ITERATIONS_HINT	Marek Olšák	2015-10-20	11	-0/+32
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	This avoids a serious r600g bug leading to a GPU hang. The chances this bug will get fixed are pretty low now. I deeply regret listening to others and not pushing this patch, leaving other users with a GPU-crashing driver. Yes, it should be fixed in the compiler and it's ugly, but users couldn't care less about that. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=86720 Cc: 11.0 10.6 <[email protected]> Reviewed-by: Brian Paul <[email protected]>
*	vc4: Switch our vertex attr lowering to being NIR-based.	Eric Anholt	2015-10-20	2	-143/+200
\| \| \| \| \| \| \| \| \| \|	This exposes more information to NIR's optimization, and should be particularly useful when we do range-based optimization. total uniforms in shared programs: 32066 -> 32065 (-0.00%) uniforms in affected programs: 21 -> 20 (-4.76%) total instructions in shared programs: 93104 -> 92630 (-0.51%) instructions in affected programs: 31901 -> 31427 (-1.49%)
*	vc4: Add limited support for ibfe/ubfe.	Eric Anholt	2015-10-20	1	-0/+42
\| \| \| \| \|	This is just enough to cover our unpack modes, which will be used by some new NIR-based lowering in the next commit.
*	radeonsi: enable BC_OPTIMIZE if centroid isn't used	Marek Olšák	2015-10-20	1	-1/+5
\| \| \| \| \| \|	This solution was recommended by a Catalyst developer. Reviewed-by: Michel Dänzer <[email protected]>
*	radeonsi: fix the export_prim_id field size in the shader key	Marek Olšák	2015-10-20	1	-2/+2
\| \| \| \|	Reviewed-by: Michel Dänzer <[email protected]>
*	radeonsi: support thread-safe shaders shared by multiple contexts	Marek Olšák	2015-10-20	9	-199/+224
\| \| \| \| \| \| \| \| \| \| \| \|	The "current" shader pointer is moved from the CSO to the context, so that the CSO is mostly immutable. The only drawback is that the "current" pointer isn't saved when unbinding a shader and it must be looked up when the shader is bound again. This is also a prerequisite for multithreaded shader compilation. Reviewed-by: Michel Dänzer <[email protected]>
*	gallium: add PIPE_CAP_SHAREABLE_SHADERS	Marek Olšák	2015-10-20	13	-0/+13
\| \| \| \| \| \|	I'll let drivers figure out how to do it. Reviewed-by: Ilia Mirkin <[email protected]>
*	radeonsi: add support for ARB_texture_view	Marek Olšák	2015-10-20	2	-7/+22
\| \| \| \| \| \| \| \| \| \| \| \|	All tests pass. We don't need to do much - just set CUBE if the view target is CUBE or CUBE_ARRAY, otherwise set the resource target. The reason this can be so simple is that texture instructions have a greater effect on the target than the sampler view. Thanks Glenn for the piglit test. Reviewed-by: Michel Dänzer <[email protected]>
*	vc4: Use nir_foreach_variable	Boyan Ding	2015-10-20	3	-7/+7
\| \| \| \| \|	Signed-off-by: Boyan Ding <[email protected]> Reviewed-by: Eric Anholt <[email protected]>
*	svga: fix incorrect round-down arithmetic	Brian Paul	2015-10-19	1	-1/+1
\| \| \| \| \| \| \| \| \|	Spotted by Roland. Luckily, this code should never really be hit since the const buffer size and offset should already be multiples of 16. I could probably add more assertions to that effect, but let's just fix the arithmetic for now. Reviewed-by: Roland Scheidegger <[email protected]>
*	ilo: set VME for 3DSTATE_PS	Chia-I Wu	2015-10-18	1	-1/+6
\| \| \| \| \|	When the bit is not set, we can see sampling artifacts on triangle edges when the mip filter is not GEN6_MIPFILTER_NONE.
*	ilo: ignore prefer_linear_threshold when zero	Chia-I Wu	2015-10-18	2	-3/+3
\| \| \| \|	This was the intended behavior but it did not work as intended until now.
*	ilo: remove some unused kernel params	Chia-I Wu	2015-10-18	2	-22/+0
\|
*	ilo: remove unused ilo_shader_get_type()	Chia-I Wu	2015-10-18	2	-12/+0
\|
*	ilo: remove u_debug.h inclusion from ilo_core.h	Chia-I Wu	2015-10-18	2	-1/+2
\| \| \| \|	Move it to ilo_debug.h.
*	ilo: remove u_memory.h inclusion from ilo_core.h	Chia-I Wu	2015-10-18	3	-1/+3
\| \| \| \|	We do not make allocations generally in the core.
*	nvc0: do not bind input params at compute state init on Fermi	Samuel Pitoiset	2015-10-18	1	-8/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	It looks like binding a constant buffer on compute overwrites the 3D state. To avoid that, we already re-bind all the 3D constant buffers after launching a compute grid but this is not enough. Binding the constant buffer of input parameters for the compute state at initialization corrupts the 3D constant buffers, and it's just useless to bind it because this is not needed until we really launch a grid. This fixes some piglit regressions related to interpolation tests introduced in "nvc0: enable compute support by default on Fermi". Fixes: 00d6186 (nvc0: enable compute support by default on Fermi) Signed-off-by: Samuel Pitoiset <[email protected]> Reviewed-by: Ilia Mirkin <[email protected]>
*	radeonsi: don't use the AMDGPU intrinsic for CMP	Marek Olšák	2015-10-17	1	-9/+22
\| \| \| \| \| \| \|	No difference according to shader-db. Reviewed-by: Michel Dänzer <[email protected]> Reviewed-by: Tom Stellard <[email protected]>
*	radeonsi: use LRP from gallivm	Marek Olšák	2015-10-17	1	-2/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Totals: SGPRS: 344552 -> 344368 (-0.05 %) VGPRS: 197132 -> 197552 (0.21 %) Code Size: 7375376 -> 7366304 (-0.12 %) bytes LDS: 91 -> 91 (0.00 %) blocks Scratch: 1679360 -> 1615872 (-3.78 %) bytes per wave Totals from affected shaders: SGPRS: 47736 -> 47552 (-0.39 %) VGPRS: 27952 -> 28372 (1.50 %) Code Size: 1392724 -> 1383652 (-0.65 %) bytes LDS: 39 -> 39 (0.00 %) blocks Scratch: 513024 -> 449536 (-12.38 %) bytes per wave Reviewed-by: Michel Dänzer <[email protected]>
*	radeonsi: don't emit AMDGPU intrinsics for integer abs, min, max	Marek Olšák	2015-10-17	1	-10/+50
\| \| \| \| \| \| \|	No difference according to shader-db. (with the new S_ABS_I32 pattern) Reviewed-by: Michel Dänzer <[email protected]> Reviewed-by: Tom Stellard <[email protected]>
*	radeonsi: don't emit AMDGPU intrinsics for EX2, ROUND, TRUNC	Marek Olšák	2015-10-17	1	-3/+3
\| \| \| \| \| \| \|	No difference according to shader-db. Reviewed-by: Michel Dänzer <[email protected]> Reviewed-by: Tom Stellard <[email protected]>
*	radeonsi: initialize output, temp, and address registers to "undef"	Marek Olšák	2015-10-17	1	-4/+15
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This removes "v_mov v0, 0" which typically occurs before exports. Totals: SGPRS: 345216 -> 344552 (-0.19 %) VGPRS: 197684 -> 197132 (-0.28 %) Code Size: 7390408 -> 7375376 (-0.20 %) bytes LDS: 91 -> 91 (0.00 %) blocks Scratch: 1842176 -> 1679360 (-8.84 %) bytes per wave Totals from affected shaders: SGPRS: 101336 -> 100672 (-0.66 %) VGPRS: 53920 -> 53368 (-1.02 %) Code Size: 2170176 -> 2155144 (-0.69 %) bytes LDS: 2 -> 2 (0.00 %) blocks Scratch: 1015808 -> 852992 (-16.03 %) bytes per wave Reviewed-by: Michel Dänzer <[email protected]> Reviewed-by: Tom Stellard <[email protected]>
*	radeonsi: implement vertex color clamping	Marek Olšák	2015-10-17	5	-4/+52
\| \| \| \| \| \|	This is only supported in the compatibility profile (without GS and tess). Reviewed-by: Michel Dänzer <[email protected]>
*	radeonsi: implement fragment color clamping	Marek Olšák	2015-10-17	6	-2/+18
\| \| \| \| \| \|	using the shader key for now. Reviewed-by: Michel Dänzer <[email protected]>
*	radeonsi: clean up other scratch buffer functions	Marek Olšák	2015-10-17	1	-15/+8
\| \| \| \|	Reviewed-by: Michel Dänzer <[email protected]>
*	radeonsi: clean up copy-pasted scratch buffer updates	Marek Olšák	2015-10-17	1	-26/+13
\| \| \| \|	Reviewed-by: Michel Dänzer <[email protected]>