I work on high performance computing and machine learning systems — LLM training and inference, GPU benchmarking, and the software that makes large models practical to run.
My objective is to affect an industry shift to massively intelligent computers that increase automation across every major industry and are super aligned to human objectives.
Reach me at gregory.diamos@gmail.com or @GregoryDiamos.
Writing
-
Outrageously Small Neural Networks: Emergent Basic Reasoning at 6,616 tok/sec on One Intel AMX Core
-
LLM Decompression: Reverse-Engineering Models Into Datasets
-
ScalarLM Benchmarking MI300X BF16 GEMM
-
ScalarLM Benchmarking MI300X Memcpy Peer
-
ScalarLM Benchmarking MI300X Memcpy
-
Tokenformer: A Scalable Transformer Architecture
-
Introducing ScalarLM v0.5: Unifying LLM Inference and Training for RL Agents
-
Building ROCm Containers for ScalarLM: A Comprehensive Guide
-
Large Language Models for Automatic Code Repair
-
Welcome to my blog
subscribe via RSS