triton

Author	SHA1	Message	Date
Philippe Tillet	4ff3714d61	[CODEGEN] Various bugfixes and stability improvements in compiler backend (#240 )	2021-08-30 11:50:35 -07:00
milesial	5b29da719d	[DRIVER] Add CUDA P2P support (#209 )	2021-08-20 21:00:54 -07:00
Philippe Tillet	298da78058	[CODEGEN/DRIVER] Tweaks for performance optimization (#193 )	2021-08-07 16:41:44 -07:00
Philippe Tillet	e8031fe61f	[DRIVER] More robust support of unsupported CUDA version (#179 )	2021-08-02 09:06:55 -07:00
Philippe Tillet	2f0f51be50	[DRIVER] No longer crashing when encountering CUDA version >11.4	2021-07-29 11:27:55 -07:00
Philippe Tillet	8eb63bcb01	[CI] Various improvements to CI (#137 ) Add clean-up before CI runs. Now using static LLVM-11 libraries from system rather than recompilation. Still no run-time LLVM dependencies	2021-07-27 12:38:49 -07:00
Philippe Tillet	94ce6aa80f	[DRIVER] Added support for CUDA 11.4 (#135 )	2021-07-27 12:38:49 -07:00
Philippe Tillet	8cea583109	[IR] Preliminary support for BF16 (#129 ) This PR adds a BF16 data-type, along with FP32 <-> BF16 conversion instructions in the LLVM codegen. Other kinds of ops on bfloat16 are not yet supported.	2021-07-27 12:38:49 -07:00
daadaada	0b05e06c0d	cu_device::max_shared_memory() now returns max dynamic shared memory size (#127 )	2021-07-27 12:38:49 -07:00
daadaada	d8d6b715c8	[CODEGEN] Performance improvement on A100 (#125 ) Improved codegen for the Ampere GPUs. * Make the layout pass recognize the multistage pipelined pattern. * Now the pipeline pass can automate the multistage pipelining transformation. * Remove extra barriers (from the prefetch pass & WAR) on Ampere. * Update the code generator (generator.cc) to make Triton generate n-buffered shared memory loads/stores.	2021-07-27 12:38:49 -07:00
Philippe Tillet	b7b05a560e	[DRIVER] Now giving the option to use system ptxas through environment variable (#123 )	2021-07-27 12:38:49 -07:00
Philippe Tillet	9f30af76fb	[GENERAL] Minor improvements: (#110 ) * Load libcuda.so.1 if libcuda.so is not there. Error if both aren't there. * Support for multiple grad_to_none in triton.testing.do_bench * Benchmark dataframe printed along with name	2021-07-27 12:38:49 -07:00
Philippe Tillet	288b4f7f58	[PYTHON] Added frontend to print sass using turingas disasm.py (#109 )	2021-07-27 12:38:49 -07:00
daadaada	967e629c0c	[CODEGEN] Add a pass to prefetch operands of dot if applicable. (#105 ) * update membar pass when data is double buffered * Add instruction prefetch_s * prefetch tests pass (except the 1 warp case) * Fix the 1-warp bug * Add back prefetch files * Disable prefetch on a100 * Always add war barrier on sm>=80	2021-07-27 12:38:49 -07:00
Philippe Tillet	1e844ba78d	[CODEGEN] Switching to predicated inline PTX for LDGs (#103 )	2021-07-27 12:38:49 -07:00
Philippe Tillet	840140bf26	[CODEGEN] Removed dedicated reassociate pass to merge it into LLVM isel (#101 ) This massively simplifies implementation of `reassociate` and also fixes a bunch of bug. The pass could still be improved, but can already be used to generate constant pointer offsets in eg the matmul epilogue	2021-07-27 12:38:49 -07:00
Philippe Tillet	39f4730305	Deprecation of Triton-C and Replacement by decorated Python functions (#86 ) This PR implements a major overhaul of the frontend for Triton, and replaces Triton-C by a pure Python API in which kernels are defined as @triton.jit decorated functions. The documentation and tutorials have also been updated to accommodate these changes. See documentations for more information on the new API	2021-07-27 12:38:49 -07:00
Philippe Tillet	5b83259592	[CODEGEN] Major performance improvements on A100 (#70 ) Improved handling of asynchronous copy, scheduling and synchronization for A100. Now achieving CUTLASS-like performance on large square dense matrix multiplication tasks	2021-07-27 12:38:49 -07:00
Philippe Tillet	3ca40b05cf	[DRIVER] Added options for developers to cache PTX file so that ti can be manually modified	2021-07-27 12:38:49 -07:00
Philippe Tillet	0b025db2ee	[RUNTIME] Added option to print LLVM-IR Also includes appropriate driver code change for that	2021-07-27 12:38:48 -07:00
Philippe Tillet	269ebc12e5	[PYTHON][TESTS][DOC] Various improvement of the API and code quality: * Simplified `triton.kernel` API to achieve lower latency: > .data_ptr() must now be passed as kernel argument. No more implicit conversion from torch.tensor > compilation options are now constant attributes, i.e., opt.d('VAR') becomes opt.VAR > torch.device must now be passed explicitly to triton.kernel (no longer inferred from torch.tensor arguments) * C++ tests moved to `python/tests/` * C++ tutorial created in `tutorials/` * Python tutorial created in python/tutorials/ * Version changed to 1.0alpha * No longer copying C++ headers into the Python package * added python/triton/ops/ package for pre-written Triton ops	2021-07-27 12:38:48 -07:00
Philippe Tillet	376c876eb8	[RUNTIME] Disable error on spills	2021-07-27 12:38:48 -07:00
Philippe Tillet	083bbd1e8d	[GENERAL] Merged v1.0alpha into master. Added features are: - A100 support via mma.16816 - Thread swizzling for conflict-free shared memory accesses without padding - Complete overhaul of the LLVM code generation in codegen/selection/generator.cc to remove overengineering - Added debugging capabilities in the Python binding - Compilation error for kernels that spill	2021-07-27 12:38:48 -07:00
Philippe Tillet	5e8f4c934c	[DRIVER] Better exception handling of invalid ptx	2021-07-27 12:38:48 -07:00
Philippe Tillet	44ca2c0cb8	[DRIVER] Removed deprecated files and functions	2021-07-27 12:38:48 -07:00
Philippe Tillet	7ab2c2a356	[DRIVER] Removed obsolete SetArg	2021-07-27 12:38:48 -07:00
Philippe Tillet	4f08d87fed	[DRIVER] Simplified Driver API by substantially removing reliance on driver::context	2021-07-27 12:38:48 -07:00
Philippe Tillet	f42b04d925	[DRIVER] Added (slow) support for CUDA11 and Ampere	2021-07-27 12:38:48 -07:00
Philippe Tillet	073fddffc1	[PYTHON] Compiling Triton in Release mode now...	2021-07-27 12:38:48 -07:00
Philippe Tillet	a77c925dfd	[DRIVER] Improved performance of Host driver code	2021-07-27 12:38:48 -07:00
Philippe Tillet	8f8d36c7a4	[GENERAL] Various bugfixes	2021-07-27 12:38:48 -07:00
Philippe Tillet	50587bbf4b	[General] LLVM-9 -> LLVM-10	2021-07-27 12:38:48 -07:00
Philippe Tillet	049ab989b5	[GENERAL] Various improvements: * Sparse einsum in triton.ops.einsum * Hacky support for fixed-tile-size atomic-add * Various bugfixes in parser	2021-07-27 12:38:48 -07:00
Philippe Tillet	664d3cae89	[DRIVER] Removed OpenCL support There is no plan to support OpenCL anytime soon (Vulkan would be preferred). Removing the adequate portion of the driver code	2021-07-27 12:38:48 -07:00
Philippe Tillet	840308ab5d	[CODEGEN] More work on the CPU backend	2021-07-27 12:38:48 -07:00
Philippe Tillet	acff1b5e05	[RUNTIME] Lower-level interface for executing functions	2021-07-27 12:38:48 -07:00
Philippe Tillet	29a0ad6c4d	[DRIVER] Now always using PTXv6.4	2021-07-27 12:38:48 -07:00
Philippe Tillet	4bb0311f60	[TRITON] Fixed misaligned address issue	2021-07-27 12:38:48 -07:00
Philippe Tillet	ddd89e1b22	[GENERAL] Fixed some undefined behavior with GCC-9	2021-07-27 12:38:48 -07:00
Philippe Tillet	3304629de9	[CORE] Fixed several issues that arose in the development of the torch-blocksparse package: * Now using warp shuffle in reductions when possible * Various bugfixes in layout inference * Added INFINITY, exponential and select * Better error messages for unimplemented constructs	2021-07-27 12:38:48 -07:00
Philippe Tillet	9cb3fd899a	[CORE][DRIVER] Now only using PTX6.4 if CUDA10.1+ is detected	2021-07-27 12:38:48 -07:00
Philippe Tillet	dfb844bf41	[GENERAL] Improved caching mechanism: * Now computing hash in libtriton * Now only compiling a single pytorch hook per function signature	2021-07-27 12:38:48 -07:00
Philippe Tillet	6d7cf35123	History prior to this date belonged to the now deprecated ISAAC project, and was deleted to save space	2021-07-27 12:38:38 -07:00

43 Commits