mesa.git - Unnamed repository; edit this file 'description' to name the repository.

	Commit message (Collapse)	Author	Age	Files	Lines
*	radeonsi/compute: Let the state tracker do all the flushing	Tom Stellard	2013-08-17	1	-3/+0
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	It shouldn't be necessary to call radeon_winsys::cs_flush() from radeonsi_launch_grid(), because the state tracker is responsible for flushing the pipeline at the appropriate time. The current behavior is also wrong, because radeonsi_launch_grid() submits packets to the compute ring, but when the state tracker calls pipe->flush() everything is submitted to the graphics ring. This has the potential to create a race condition. The downside of removing this flush is that the compute dispatch packets will be sent to the graphics ring rather than the compute ring. In the future we will need to come up with a way to detect 'compute' command streams and submit them to the appropriate ring. Signed-off-by: Marek Olšák <[email protected]>
*	i965: Dump more information about batch buffer usage.	Kenneth Graunke	2013-08-16	1	-3/+10
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Previously, INTEL_DEBUG=bat would dump messages like: intel_mipmap_tree.c:1643: Batchbuffer flush with 456b used This only reported the space used for command packets, and didn't report any information on the space used for indirect state. Now it dumps: intel_context.c:366: Batchbuffer flush with 6128b (pkt) + 4288b (state) = 10416b (31.8%) This conveniently shows the breakdown of space used for packets vs. state, as well as the percentage of batchbuffer space. Signed-off-by: Kenneth Graunke <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965: Add Gen7 depth stall flushes before disabling depth in BLORP.	Kenneth Graunke	2013-08-16	1	-0/+2
\| \| \| \| \| \| \| \| \| \|	We emit these before configuring depth in the normal path, or actually using the depth buffer in BLORP - we just failed to emit them when disabling depth altogether. Signed-off-by: Kenneth Graunke <[email protected]> Reviewed-by: Chad Versace <[email protected]> Reviewed-by: Ian Romanick <[email protected]>
*	i965: Add Gen6 depth stall flushes before disabling depth in BLORP.	Kenneth Graunke	2013-08-16	1	-0/+3
\| \| \| \| \| \| \| \| \| \| \| \|	We emit these before configuring depth in the normal path, or actually using the depth buffer in BLORP - we just failed to emit them when disabling depth altogether. On Sandybridge, this also requires the post_sync_nonzero flush. Signed-off-by: Kenneth Graunke <[email protected]> Reviewed-by: Chad Versace <[email protected]> Reviewed-by: Ian Romanick <[email protected]>
*	i965: Don't copy propagate bitcasts with source modifiers.	Matt Turner	2013-08-16	3	-4/+23
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Previously, copy propagation would cause bitcast_f2u(abs(float)) to be performed in a single step, but the application of source modifiers (abs, neg) happens after type conversion, leading to incorrect results. That is, for bitcast_f2u(abs(float)) we would in fact generate code to do abs(bitcast_f2u(float)). For example, whereas bitcast_f2u(abs(float)) might result in a register argument such as (abs)g2.2<0,1,0>UD v2: Set interfered = true and break in register_coalesce instead of returning false. Reviewed-by: Paul Berry <[email protected]>
*	i965: Emit MOVs for neg/abs.	Matt Turner	2013-08-16	2	-4/+4
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Necessary to avoid combining a bitcast and a modifier into a single operation. Otherwise if safe, the MOV should be removed by copy-propagation or register coalescing. With this and the next patch, there are only four changes in shader-db: all a single extra instruction. The code does something like mov a.w, -b.x and copy propagation doesn't work because it only handles no-op swizzles. Seems acceptable, given the known limitation of our copy propagation. Reviewed-by: Ian Romanick <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/blorp: Add support for single sample scaled blit with bilinear filter	Anuj Phogat	2013-08-16	3	-38/+69
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Currently single sample scaled blits with GL_LINEAR filter falls back to meta path. Patch removes this limitation in BLORP engine and implements single sample scaled blit with bilinear filter. No piglit, gles3 regressions are observed with this patch on Ivybridge. V2: Use "sample" message to utilize the linear filtering functionality built in to hardware. V3: Define a bool variable (bilinear_filter) to handle the conditions for GL_LINEAR blits. Signed-off-by: Anuj Phogat <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/blorp: Define a function to clamp texture coordinates	Anuj Phogat	2013-08-16	1	-24/+39
\| \| \| \| \| \| \| \| \| \|	New function clamp_tex_coords() clamps the texture coordinates to texture boundaries. This function will also be utilized later for the BLORP implementation of single-sample scaled blit with bilinear filter. Signed-off-by: Anuj Phogat <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/blorp: Use more appropriate variable names	Anuj Phogat	2013-08-16	2	-18/+14
\| \| \| \| \| \| \| \| \| \| \|	When we talk about both multi-sample and single-sample scaled blits, rect_grid_{x1, y1} are more appropriate variable names as compared to sample_grid_{x1, y1}. There are no functional changes in this patch. It just prepares for the BLORP implementation of single-sample scaled blit with bilinear filter. Signed-off-by: Anuj Phogat <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	meta: Fix blitting a framebuffer with renderbuffer attachment	Anuj Phogat	2013-08-16	1	-10/+15
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch fixes a case of framebuffer blitting with renderbuffer as color attachment and GL_LINEAR filter. Meta implementation of glBlitFrambuffer() converts source color buffer to a texture and uses it to do the scaled blitting in to destination buffer. Using the exact source rectangle to create the texture does incorrect linear filtering along the edges. This patch makes the changes to extend the texture edges by one pixel in x, y directions. This ensures correct linear filtering. It fixes failing piglit fbo-attachments-blit-scaled-linear test. Signed-off-by: Anuj Phogat <[email protected]> CC: "9.2" <[email protected]> CC: "9.1" <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	nv50: add vp3/vp4 support for mpeg2/vc1	Ilia Mirkin	2013-08-16	12	-12/+927
\| \| \| \| \| \| \|	h264/mpeg4 remain disabled for pre-nvc0, there's some minor bug/difference which causes the decoding to hang after some frames. Signed-off-by: Ilia Mirkin <[email protected]>
*	nv50: separate video logic from noalloc	Ilia Mirkin	2013-08-16	3	-3/+6
\| \| \| \| \| \| \|	The upcoming vp3 logic will want the video layout, but allocated by the miptree. Signed-off-by: Ilia Mirkin <[email protected]>
*	nv30: remove no-longer-used formats from table	Ilia Mirkin	2013-08-16	1	-3/+0
\| \| \| \| \| \| \| \|	Commit 14ee790df77 removed the formats from the vtxfmt_table but forgot to also update the info_table. Signed-off-by: Ilia Mirkin <[email protected]> Cc: "9.2 and 9.1" <[email protected]>
*	mesa: Update the BGRA vertex array error handling	Fredrik Höglund	2013-08-15	1	-1/+19
\| \| \| \| \| \| \| \|	The error code was changed from INVALID_VALUE to INVALID_OPERATION in OpenGL 3.3. We should also generate an error when size is BGRA and normalized is FALSE. Reviewed-by: Kenneth Graunke <[email protected]>
*	i965/fs: Fix Sandybridge regressions from SEL optimization.	Kenneth Graunke	2013-08-15	1	-4/+13
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Sandybridge is the only platform that supports an IF instruction with an embedded comparison. In this case, we need to emit a CMP to go along with the SEL. Fixes regressions in Piglit's glsl-fs-atan-3, fs-unpackHalf2x16, fs-faceforward-float-float-float, isinf-and-isnan fs_basic, and isinf-and-isnan fs_fbo. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=68086 Signed-off-by: Kenneth Graunke <[email protected]> Reviewed-by: Matt Turner <[email protected]> Reviewed-by: Anuj Phogat <[email protected]> Tested-by: lu hua <[email protected]>
*	i965: Force X-tiling for 128 bpp formats on Sandybridge.	Kenneth Graunke	2013-08-15	1	-0/+9
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	128 bpp formats are not allowed to be Y-tiled on any architectures except Gen7. +11 Piglits on Sandybridge (mostly regression fixes since the switch to Y-tiling). Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=63867 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=64261 Signed-off-by: Kenneth Graunke <[email protected]> Reviewed-by: Chad Versace <[email protected]> Reviewed-by: Ian Romanick <[email protected]> Cc: "9.2" <[email protected]>
*	mesa/vbo: Fix handling of attribute 0 in non-compatibilty contexts	Ian Romanick	2013-08-15	1	-23/+59
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	It is only in OpenGL compatibility-style contexts where generic attribute 0 and GL_VERTEX_ARRAY have a bizzare, aliasing relationship. Moreover, it is only in OpenGL compatibility-style contexts and OpenGL ES 1.x where one of these attributes provokes the vertex. In all other APIs each implicit call to glArrayElement provokes a vertex regardless of which attributes are enabled. Signed-off-by: Ian Romanick <[email protected]> Reviewed-by: Robert Bragg <[email protected]> Cc: "9.0 9.1 9.2" <[email protected]> Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=55503 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=66292 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=67548
*	draw: handle nan clipdistance	Zack Rusin	2013-08-15	5	-4/+48
\| \| \| \| \| \| \| \|	If clipdistance for one of the vertices is nan (or inf) then the entire primitive should be discarded. Signed-off-by: Zack Rusin <[email protected]> Reviewed-by: Roland Scheidegger <[email protected]>
*	i915,i965: Fix memory leak in try_pbo_upload (v2)	Vinson Lee	2013-08-15	2	-0/+2
\| \| \| \| \| \| \| \| \| \| \|	Fixes "Resource leak" defect reported by Coverity. Tested on Haswell, no Piglit regressions. v2: Apply to i965, not just i915. (chadv) CC: "9.2, 9.1" <[email protected]> Signed-off-by: Vinson Lee <[email protected]> Reviewed-by: Chad Versace <[email protected]>
*	gallivm: revert accidentally commited hunk	Roland Scheidegger	2013-08-15	1	-12/+1
\| \| \| \|	That magic wasn't meant to be commited, need to work on some proper fix.
*	gallivm: do per-sample depth comparison instead of doing it post-filter	Roland Scheidegger	2013-08-15	2	-106/+195
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Doing the comparisons pre-filter is highly recommended by OpenGL (and d3d9) and definitely required by d3d10. This actually doesn't do it pre-filter but more "in-filter" as otherwise need to push the comparisons even further down into fetch code and this also trivially allows using a somewhat cheaper lerp. Doing it pre-filter would actually have some performance advantage for UNORM formats (because the comparisons should be done in texture format, we'd only need to convert the shadow ref coord to texture format once, but in turn would save converting the per-sample texture values to floats) but this gets a bit messy as this has implications for border color handling as well (which needs to be done prior to depth comparisons, hence would also need to convert border color to texture format too or use some other tricks like doing separate border color / shadow ref comparison and simply using that result directly when doing border replacement). Should make no difference for nearest filtering, and performance for linear filtering should be mostly the same too (essentially have one more comparison instruction per sample, and replace the sub/mul/add lerp with a sub/and/and/add special "lerp" which all in all shouldn't be much of a difference). v2: get rid of old code completely Reviewed-by: Zack Rusin <[email protected]>
*	radeonsi: Pixel shaders pre-load one more SGPR	Michel Dänzer	2013-08-15	1	-2/+3
\| \| \| \|	Acked-by: Marek Olšák <[email protected]>
*	radeonsi: TGSI_SEMANTIC_CLIPVERTEX doesn't use any parameters	Michel Dänzer	2013-08-15	1	-0/+1
\|
*	radeonsi: Don't export unused clip distance vectors from vertex shader	Michel Dänzer	2013-08-15	3	-1/+14
\| \| \| \| \| \| \| \|	E.g. the Source engine seems to always write to gl_ClipVertex, but normally doesn't enable any GL_CLIP_DISTANCEn states. This change removes some irrelevant parts from the generated vertex shader code in such cases. Reviewed-by: Tom Stellard <[email protected]>
*	radeonsi: Don't leave gaps between position exports from vertex shader	Michel Dänzer	2013-08-15	3	-59/+83
\| \| \| \| \| \| \| \| \| \| \|	If the vertex shader exports clip distances but not point size, use position exports 1/2 instead of 2/3 for the clip distances. Fixes geometry corruption in that case. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=66974 Cc: [email protected] Reviewed-by: Tom Stellard <[email protected]>
*	llvmpipe: fix stencil bug if we have both stencil and depth tests	Roland Scheidegger	2013-08-15	1	-14/+13
\| \| \| \| \| \| \| \| \| \| \| \| \|	This is a very well hidden bug found by accident (only the fixed glean tstencil2 test so far seems to hit it). We must use new mask with combined s_pass values and orig_mask values for zpass/zfail stencil ops, otherwise both the sfail op and one of zpass/zfail op are applied (probably not hit in most tests because some of the ops tend to be KEEP usually). Note: this is a candidate for the 9.2 branch. Reviewed-by: Zack Rusin <[email protected]>
*	st/mesa: use new float comparison opcodes if native integers are supported	Roland Scheidegger	2013-08-15	1	-34/+26
\| \| \| \| \| \| \| \| \|	Should get rid of some float-to-int conversions (with negation). No piglit regressions (with llvmpipe). v2: fix bogus formatting spotted by Brian. Reviewed-by: Brian Paul <[email protected]>
*	nvc0: move video param and format support functions to nouveau	Ilia Mirkin	2013-08-15	5	-70/+76
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: move firmware loading functions to nouveau	Ilia Mirkin	2013-08-15	3	-90/+108
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: move some of the simpler decoder functions into nouveau	Ilia Mirkin	2013-08-15	3	-62/+69
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: move vp param filling logic into nouveau	Ilia Mirkin	2013-08-15	6	-476/+499
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: move bsp param-filling logic into nouveau	Ilia Mirkin	2013-08-15	4	-276/+324
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: move nvc0_decoder into nouveau, rename to nouveau_vp3_decoder	Ilia Mirkin	2013-08-15	6	-224/+227
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: standardize on using #if for NVC0_DEBUG_FENCE	Ilia Mirkin	2013-08-15	5	-8/+8
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: refactor video buffer management logic into nouveau_vp3	Ilia Mirkin	2013-08-15	8	-175/+243
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nv50: allow forcing PMPEG use, for ease of testing	Ilia Mirkin	2013-08-15	2	-2/+4
\| \| \| \| \| \| \|	This also allows people who don't want to install the binary blobs required for VP2 to still get MPEG decoding. Signed-off-by: Ilia Mirkin <[email protected]>
*	nv30: hook up PMPEG support via nouveau_video, enables XvMC to work	Ilia Mirkin	2013-08-15	3	-15/+15
\| \| \| \| \| \| \|	Force the format to be the reasonable format that doesn't require an inverse z-scan. Signed-off-by: Ilia Mirkin <[email protected]>
*	nouveau: set buffer format of video buffer	Ilia Mirkin	2013-08-15	1	-0/+1
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nouveau: fix number of surfaces in video buffer, use defines	Ilia Mirkin	2013-08-15	1	-4/+4
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nv30: U8_USCALED only works for size 4	Ilia Mirkin	2013-08-15	1	-3/+0
\| \| \| \| \| \| \| \| \|	See https://bugs.freedesktop.org/show_bug.cgi?id=61635 for a sample program. Changing it to use a vec4 makes it work. Remove the unsupported formats. Signed-off-by: Ilia Mirkin <[email protected]> Cc: "9.2 and 9.1" <[email protected]>
*	i965: allow 8 user clip planes on CTG+	Chris Forbes	2013-08-16	5	-6/+21
\| \| \| \| \| \| \| \| \| \| \| \| \|	There's no need to use a clip flag for NEGW on these gens, so no reason we can't just enable 8 planes. V2: - Bump (and document!) MAX_VERTS in the clip code. - Fix clip flag masks in the clip unit state and in the shader prolog - Move this to the end of the series for less breakage. Signed-off-by: Chris Forbes <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965: get rid of clip plane compaction	Chris Forbes	2013-08-16	4	-55/+11
\| \| \| \| \|	Signed-off-by: Chris Forbes <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/clip: Support clip distances for line clipping	Chris Forbes	2013-08-16	1	-19/+47
\| \| \| \| \| \| \| \| \|	This does the same thing as we do for triangle clipping -- select the appropriate source (either dot(hpos,fixed plane) or a clipdistance slot). Signed-off-by: Chris Forbes <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/clip: remove spurious clipvertex param	Chris Forbes	2013-08-16	1	-11/+4
\| \| \| \| \| \| \| \| \| \|	Nothing in the clipper uses gl_ClipVertex any more, so we don't care where it is. V2: Don't bother fishing out the clipvertex offset either. Signed-off-by: Chris Forbes <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/clip: Use clip distances for all user clipping	Chris Forbes	2013-08-16	1	-5/+6
\| \| \| \| \| \| \|	V2: Adjust explanation of load_clip_distance() Signed-off-by: Chris Forbes <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i956/clip: push dp4 into load_clip_distance	Chris Forbes	2013-08-16	1	-17/+22
\| \| \| \| \| \| \| \| \| \|	Soon the dp4 is only going to be used for fixed clip planes. V2: Remove old inaccurate comment about the behavior of this function; add a better explanation above. Signed-off-by: Chris Forbes <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/clip: Track offset into the vertex for clipdistance	Chris Forbes	2013-08-16	2	-0/+12
\| \| \| \| \|	Signed-off-by: Chris Forbes <[email protected]> Reviewed-by: Paul Berry <[email protected]>
*	i965/Gen4-5: Set clip flags from clip distances	Chris Forbes	2013-08-16	1	-11/+11
\| \| \| \| \| \| \| \| \| \| \| \| \|	V2: - Use the new VS_OPCODE_UNPACK_FLAGS_SIMD4X2 to correctly split the flags for the two vertices being processed together. - Don't apply bogus masking of clip flags. The set of plane enables aren't included in the shader key, and we wouldn't want the recompiles anyway. V3: - Tidy up spurious instructions, name temps properly. Signed-off-by: Chris Forbes <[email protected]> [V2] Reviewed-by: Paul Berry <[email protected]>
*	i965: add new VS_OPCODE_UNPACK_FLAGS_SIMD4X2	Chris Forbes	2013-08-16	4	-1/+29
\| \| \| \| \| \| \| \| \| \| \| \| \| \|	Splits the bottom 8 bits of f0.0 for further wrangling in a SIMD4x2 program. The 4 bits corresponding to the channels in each program flow are copied to the LSBs of dst.x visible to each flow. This is useful for working with clipping flags in the VS. V3: - Fixup immediate types - Teach scheduler about the hidden dep on flags Signed-off-by: Chris Forbes <[email protected]> V2: Reviewed-by: Paul Berry <[email protected]>
*	i965/vs: add vec4_instruction::depends_on_flags	Chris Forbes	2013-08-16	2	-2/+7
\| \| \| \| \| \| \|	We're about to have an instruction that depends on the flags but isn't predicated. This lays the groundwork. Signed-off-by: Chris Forbes <[email protected]>