AMD Launches Triton Inference Server with ONNX Runtime Backend
AMD released Triton Inference Server with ONNX Runtime backend for AMD GPUs, supporting dynamic batching and MIGraphX optimization.
Browse all published articles.
AMD released Triton Inference Server with ONNX Runtime backend for AMD GPUs, supporting dynamic batching and MIGraphX optimization.
AMD's Gluon GEMM tutorial on MI355 GPUs reaches 99% MFMA efficiency with FP16 kernels, achieving 1489 TFLOPS through iterative optimization.
AMD Ryzen AI Max+ processors with 128 GB unified memory allow running 100B+ parameter models locally, eliminating the need for multiple GPUs or cloud services.
Alta Daily, a fashion app, uses Meta's Segment Anything Model to process over 20 million images, helping users visualize outfits on their avatars.
Meta released Segment Anything Model 3 (SAM 3.1), improving video processing speed by doubling throughput to 32 frames per second using a single H100 GPU.
Meta has deployed hundreds of thousands of MTIA chips to power AI experiences for billions of users, with new generations set for 2026 and 2027.
AMD's 4-wave interleave FP8 GEMM design improves performance by doubling register file size, enabling full 128×128 output tile storage. The update builds on prior 8-wave ping-pong implementation.
Hugging Face's server capacity reached 363 sessions, exceeding the 200-session limit, causing disruptions for users since over two days.
xAI announced today that users with SuperGrok or X Premium+ subscriptions can now use Grok models inside Kilo Code, an open-source agentic coding platform.
X announced Grok Build, a new coding agent available in early beta for SuperGrok and X Premium Plus users, on May 25, 2026.
xAI enables Grok model usage in OpenClaw, an open-source personal assistant, starting May 19, 2026.
xAI today launched Grok Skills, a new feature that allows users to teach Grok once and have it remember across all conversations, with no setup required.