Hardware

891 articles found

New GitHub Workshop Teaches Engineers to Run AI Agents on Dedicated GPUs Without Third-Party APIs

New GitHub Workshop Teaches Engineers to Run AI Agents on Dedicated GPUs Without Third-Party APIs

Aug 18, 2026
GitHub

A new GitHub workshop, 'Agents That Own Their Inference,' teaches engineers to build production AI agents on dedicated GPUs using vLLM, covering nine modules on inference optimization and Kubernetes deployment with full open-source control and no vendor lock-in, available for roughly $0.52 per hour on Akamai Cloud.

AI Agent Costs Set to Surge Fivefold by 2028 as Enterprises Race to Adopt Despite Hidden Expenses

AI Agent Costs Set to Surge Fivefold by 2028 as Enterprises Race to Adopt Despite Hidden Expenses

Aug 18, 2026
The Deep View

AI agent costs are set to surge fivefold by 2028 as enterprises rush to adopt the technology despite hidden expenses and unpredictable returns, with Gartner warning of an 'inference paradox' where efficiency gains actually drive higher spending, while Nvidia's $105 billion financing of OpenAI's Ohio data center highlights just how …

Alibaba's New 27B Vision AI Model Runs on Laptops But Overthinks Simple Tasks

Alibaba's New 27B Vision AI Model Runs on Laptops But Overthinks Simple Tasks

Aug 17, 2026
Simon Willison’s Weblog

Alibaba's new Qwen 3.8 27B vision AI model runs on consumer laptops as a 17GB file, delivering impressive code generation and image analysis capabilities, but its default 'xhigh' reasoning setting causes excessive overthinking on simple tasks, and speed remains limited at 15–30 tokens per second — though a 72% boost …

New Open-Source Tool 'Soup' Lets Developers Fine-Tune 8B AI Models on a 4GB Laptop GPU With a Single Command

New Open-Source Tool 'Soup' Lets Developers Fine-Tune 8B AI Models on a 4GB Laptop GPU With a Single Command

Aug 17, 2026
GitHub

A new open-source tool called 'Soup' is revolutionizing AI development by allowing developers to fine-tune massive 8-billion parameter language models on a standard 4GB laptop GPU using a single command, thanks to a breakthrough 'layer streaming' technique that achieves 119.6 tokens per second at just 3.32 GB peak VRAM.

Page 1 of 90
Next
Showing 1 - 10 of 891 articles