mesa.git - Unnamed repository; edit this file 'description' to name the repository.

	Commit message (Collapse)	Author	Age	Files	Lines
*	freedreno/ir3: split out shader compiler from a3xx	Rob Clark	2014-07-25	25	-477/+580
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Move the bits we want to share between generations from fd3_program to ir3_shader. So overall structure is: fdN_shader_stateobj -> ir3_shader -> ir3_shader_variant -> ir3 \|- ... \- ir3_shader_variant -> ir3 So the ir3_shader becomes the topmost generation neutral object, which manages the set of variants each of which generates, compiles, and assembles it's own ir. There is a bit of additional renaming to s/fd3_compiler/ir3_compiler/, etc. Keep the split between the gallium level stateobj and the shader helper object because it might be a good idea to pre-compute some generation specific register values (ie. anything that is independent of linking). Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx/compiler: rename ir3_shader to ir3	Rob Clark	2014-07-25	12	-55/+55
\| \| \| \| \| \| \| \|	First step of reoganization split out compiler (so it can be shared between a3xx and a4xx). Rename ir3_shader -> ir3 (since we'll want the name ir3_shader for a higher level object). Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx/compiler: scheduler vs pred reg	Rob Clark	2014-07-25	2	-3/+51
\| \| \| \| \| \| \|	The scheduler also needs to be aware of predicate register (p0) in addition to address register (a0). Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx/compiler: little cleanups	Rob Clark	2014-07-25	4	-39/+19
\| \| \| \| \| \|	Remove some obsolete comments, rename deref->addr. Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx: enable/disable wa's based on patch-level	Rob Clark	2014-07-25	4	-8/+34
\| \| \| \| \| \| \| \|	It seems like for the most part, different behaviors, workarounds, etc, should be conditional on GPU patch revision (ie. a320.0 vs a320.2) rather than GPU id (a320 vs a330). Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx/compiler: make IR heap dyanmic	Rob Clark	2014-07-25	2	-8/+43
\| \| \| \| \| \| \| \| \| \| \| \|	The fixed size heap is a remnant of the fdre-a3xx assembler. Yet it is convenient for being able to free the entire data structure in one shot without worrying about leaking nodes. Change it to dynamically grow the heap size (adding chunks) as needed so we don't have an artificial upper limit on shader size (other than hw limits) and don't always have to allocate worst-case size. Signed-off-by: Rob Clark <[email protected]>
*	r600g/compute: Fix singed/unsigned comparison compiler warnings.	Jan Vesely	2014-07-25	1	-7/+7
\| \| \| \| \| \| \|	The iteration variables go from 0 anyway. Signed-off-by: Jan Vesely <[email protected]> Reviewed-by: Tom Stellard <[email protected]>
*	clover: Query the device to see if images are supported	Tom Stellard	2014-07-25	3	-1/+8
\| \| \| \|	Reviewed-by: Francisco Jerez <[email protected]>
*	gallium: Add PIPE_CAP_COMPUTE_IMAGES_SUPPORTED	Tom Stellard	2014-07-25	3	-1/+11
\| \| \| \| \|	Reviewed-by: Marek Olšák <[email protected]> Reviewed-by: Francisco Jerez <[email protected]>
*	r600g/compute: Allow compute_memory_defrag to defragment between resources	Bruno Jiménez	2014-07-25	2	-5/+7
\| \| \| \| \| \|	This will be used in the following patch to avoid duplicated code Reviewed-by: Tom Stellard <[email protected]>
*	r600g/compute: Allow compute_memory_move_item to move items between resources	Bruno Jiménez	2014-07-25	2	-16/+16
\| \| \| \| \| \|	v2: Remove unnecesary variables Reviewed-by: Tom Stellard <[email protected]>
*	winsys/radeon: fix indentation	Jerome Glisse	2014-07-24	1	-29/+29
\| \| \| \| \| \| \|	Can we please keep it clean and avoid ending up in messy situation like ddx. Signed-off-by: Jérôme Glisse <[email protected]>
*	nvc0/ir: support 2d constbuf indexing	Ilia Mirkin	2014-07-24	1	-0/+14
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	gm107/ir: emit LDC subops	Ilia Mirkin	2014-07-24	1	-0/+1
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	gk110/ir: emit load constant subop	Ilia Mirkin	2014-07-24	1	-0/+1
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nv50/ir: fix phi/union sources when their def has been merged	Ilia Mirkin	2014-07-24	1	-0/+8
\| \| \| \| \| \| \| \| \| \| \| \| \|	In a situation where double-register values are used, the phi nodes can still end up being u32 values. They all get merged into one RA node though. When fixing up the merge (which comes after the phi node), the phi node's def would get fixed, but not its sources which would remain at the low register value. This maintains the invariant that a phi node's defs and sources are allocated the same register. Signed-off-by: Ilia Mirkin <[email protected]>
*	nv50/ir: fix hard-coded TYPE_U32 sized register	Ilia Mirkin	2014-07-24	1	-3/+4
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: mark shader header if fp64 is used	Ilia Mirkin	2014-07-24	1	-0/+2
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nv50/ir: keep track of whether the program uses fp64	Ilia Mirkin	2014-07-24	2	-2/+7
\| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]>
*	nvc0: make sure that the local memory allocation is aligned to 0x10	Ilia Mirkin	2014-07-24	1	-1/+1
\| \| \| \| \|	Signed-off-by: Ilia Mirkin <[email protected]> Cc: <[email protected]>
*	ilo: check the tilings of imported handles	Chia-I Wu	2014-07-24	1	-30/+36
\| \| \| \|	Just to be cautious.
*	ilo: clean up resource bo renaming	Chia-I Wu	2014-07-24	4	-51/+63
\| \| \| \| \|	s/alloc_bo/rename_bo/ as that is what the functions do. Simplify bo allocation and move the complexity to bo renaming.
*	ilo: share some code between {tex,buf}_create_bo	Chia-I Wu	2014-07-24	1	-59/+55
\| \| \| \| \|	Add resource_get_bo_name() and resource_get_bo_initial_domain() for use by both functions.
*	ilo: use native 3-component vertex formats on GEN7.5+	Chia-I Wu	2014-07-24	2	-1/+6
\| \| \| \|	GEN7.5 gains support for those formats natively.
*	ilo: allow for device-dependent format translation	Chia-I Wu	2014-07-24	5	-32/+39
\| \| \| \|	Pass ilo_dev_info to all format translation functions.
*	freedreno/a3xx/compiler: fix p0 (kill, etc)	Rob Clark	2014-07-23	1	-1/+2
\| \| \| \| \| \| \|	Don't assert (debug builds) or assign random uninitialized value for predicate register (p0).. that screws up kill, etc. Signed-off-by: Rob Clark <[email protected]>
*	Revert "r600g/compute: Fix warnings"	Tom Stellard	2014-07-23	2	-16/+12
\| \| \| \| \| \|	This reverts commit 467f1585e28adba0e94ef593de131bc327f098bb. This breaks the build on some systems.
*	radeon/llvm: fix formatting	Grigori Goronzy	2014-07-23	1	-10/+14
\| \| \| \| \| \| \|	Use K&R and same indent as most other code. No functional change intended. Reviewed-by: Tom Stellard <[email protected]>
*	radeon/llvm: enable unsafe math for graphics shaders	Grigori Goronzy	2014-07-23	1	-0/+5
\| \| \| \| \| \| \| \| \| \| \|	Accuracy of some operations was recently improved in the R600 backend, at the cost of slower code. This is required for compute shaders, but not for graphics shaders. Add unsafe-fp-math hint to make LLVM generate faster but possibly less accurate code. Piglit didn't indicate any regressions. Reviewed-by: Tom Stellard <[email protected]>
*	r600g/compute: Fix warnings	Tom Stellard	2014-07-23	2	-12/+16
\|
*	r600g: Use hardware sqrt instruction	Glenn Kennard	2014-07-23	2	-7/+4
\| \| \| \| \| \| \|	Piglit quick tests including sqrt pass, no other regressions, tested on radeon 6670. Reviewed-by: Alex Deucher <[email protected]>
*	r600g/compute: Remove unneeded code from compute_memory_promote_item	Bruno Jiménez	2014-07-23	2	-36/+12
\| \| \| \| \| \| \| \| \| \| \| \|	Now that we know that the pool is defragmented, we positively know that allocated + unallocated will be the total size of the current pool plus all the items that will be promoted. So we only need to grow the pool once. This will allow us to just add the new items to the end of the item_list without the need of looking for a place to the new item. Reviewed-by: Tom Stellard <[email protected]>
*	r600g/compute: Quick exit if there's nothing to add to the pool	Bruno Jiménez	2014-07-23	1	-0/+4
\| \| \| \| \| \| \| \|	This way we can avoid defragmenting the pool, even if it is needed to defragment it, and looping again through the list of unallocated items. Reviewed-by: Tom Stellard <[email protected]>
*	r600g/compute: Defrag the pool if it's necesary	Bruno Jiménez	2014-07-23	2	-17/+19
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This patch adds a new member to the pool to track its status. For now it is used only for the 'fragmented' status, but if needed it could be used for more statuses. The pool will be considered fragmented if: An item that isn't the last is freed or demoted. This 'strategy' has a problem, although it shouldn't cause any bug. If for example we have two items, A and B. We choose to free A first, now the pool will have the 'fragmented' status. If we now free B, the pool will retain its 'fragmented' status even if it isn't fragmented. Reviewed-by: Tom Stellard <[email protected]>
*	r600g/compute: Add a function for defragmenting the pool	Bruno Jiménez	2014-07-23	2	-0/+28
\| \| \| \| \| \| \| \| \| \| \|	This new function will move items forward in the pool, so that there's no gap between them, effectively defragmenting the pool. For now this function is a bit dumb as it just moves items forward without trying to see if other items in the pool could fit in the gaps. Reviewed-by: Tom Stellard <[email protected]>
*	r600g/compute: Add a function for moving items in the pool	Bruno Jiménez	2014-07-23	2	-0/+93
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	This function will be used in the future by compute_memory_defrag to move items forward in the pool. It does so by first checking for overlaping ranges, if the ranges don't overlap it will copy the contents directly. If they overlap it will try first to make a temporary buffer, if this buffer fails to allocate, it will finally fall back to a mapping. Note that it will only be needed to move items forward, it only checks for overlapping ranges in that case. If needed, it can easily be added by changing the first if. Reviewed-by: Tom Stellard <[email protected]>
*	freedreno/a3xx: more vtx formats	Rob Clark	2014-07-23	1	-0/+17
\| \| \| \| \| \| \| \|	Actually what we currently handle is just the SCALED versions, and not the int versions. The difference probably matters more when we actually support integer in the compiler. Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx/compiler: const file relative addressing	Rob Clark	2014-07-23	8	-68/+203
\| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \| \|	Teach new compiler scheduling and register assignment how to deal with relative addressing. This gets us what we need to avoid falling back to old compiler for CONST[ADDR[0].x+n]. It is also a prerequisite for temp file relative addressing, although that is going to also need some cleverness in register assignment to keep arrays grouped together. NOTE: doing address calculation in full precision and then narrowing to s16 in the mov to addr reg seems to sometimes cause lockups (and sometimes work?!). It seems more reliable to do the address calculation in s16, like the blob does. Which means teaching RA how to deal with mixed half and full precision allocation. Fortunately that didn't turn out to be too hard, so that is a nice bonus which we could probably take better advantage of elsewhere. Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx/compiler: move function	Rob Clark	2014-07-23	1	-35/+35
\| \| \| \|	Signed-off-by: Rob Clark <[email protected]>
*	freedreno/a3xx: add back a few stalls	Rob Clark	2014-07-23	1	-0/+8
\| \| \| \| \| \| \| \| \| \| \|	Technically we should not need these. CP_LOAD_STATE can be pipelined. But removing them broke a few piglit tests, like fbo-depth- GL_DEPTH_COMPONENT24-readpixels. I expect these are just masking a problem elsewhere, or perhaps they are only needed under some more specific circumstances. But until that is understood properly, give back a bit of the perf boost we got from c63450e8. Signed-off-by: Rob Clark <[email protected]>
*	targets/dri: fix freedreno targets	Rob Clark	2014-07-23	2	-3/+11
\| \| \| \| \| \| \|	The kernel driver name is either "kgsl" (downstream/android) or "msm" (upstream). Signed-off-by: Rob Clark <[email protected]>
*	freedreno: update generated headers	Rob Clark	2014-07-23	4	-14/+14
\| \| \| \|	Signed-off-by: Rob Clark <[email protected]>
*	r600g/radeonsi: Use write-combined CPU mappings of some BOs in GTT	Michel Dänzer	2014-07-23	17	-26/+77
\| \| \| \|	Reviewed-by: Marek Olšák <[email protected]>
*	winsys/radeon: Use separate caching buffer managers for VRAM and GTT	Michel Dänzer	2014-07-23	3	-9/+20
\| \| \| \| \| \| \|	Should reduce overhead because the caching buffer manager doesn't need to consider buffers of the wrong type. Reviewed-by: Marek Olšák <[email protected]>
*	radeonsi/compute: Add support scratch buffer support v2	Tom Stellard	2014-07-21	3	-2/+85
\| \| \| \| \| \| \| \|	The scratch buffer will be used for private memory and also register spilling. v2: - Code cleanups
*	radeonsi/compute: Bump number of user sgprs for LLVM 3.5	Tom Stellard	2014-07-21	1	-1/+6
\| \| \| \|	Reviewed-by: Marek Olšák <[email protected]>
*	winsys/radeon: Query the kernel for the number of SEs and SHs per SE	Tom Stellard	2014-07-21	2	-0/+8
\| \| \| \|	Reviewed-by: Marek Olšák <[email protected]>
*	radeonsi/compute: Share COMPUTE_DBG macro with r600g	Tom Stellard	2014-07-21	3	-13/+10
\| \| \| \|	Reviewed-by: Marek Olšák <[email protected]>
*	radeonsi: Read rodata from ELF and append it to the end of shaders	Tom Stellard	2014-07-21	3	-1/+22
\| \| \| \| \| \| \|	The is used for programs that have arrays of constants that are accessed using dynamic indices. The shader will compute the base address of the constants and then access them using SMRD instructions.
*	radeonsi: only update vertex buffers when they need updating	Marek Olšák	2014-07-18	3	-2/+22
\| \| \| \|	Reviewed-by: Michel Dänzer <[email protected]>