Everything you care about in one place

Follow feeds: blogs, news, RSS and more. An effortless way to read and digest content of your choice.

Get Feeder

digitalocean.com

DigitalOcean Community Tutorials

Get the latest updates from DigitalOcean Community Tutorials directly as they happen.

Follow now 91 followers

Latest posts

Last updated about 2 hours ago

Making an AI Agent Feel Fast on Serverless Inference

about 4 hours ago

Every team building an agent has access to the same models. Very...

Why Spiky Inference Traffic Breaks the Dedicated GPU Math

1 day ago

Introduction A dedicated GPU running spiky LLM inference traffic clears a specific,...

Where to Run Qwen 3 in Production: Inference Providers Compared

2 days ago

Selecting a Qwen model is only the first production decision. The next...

Multi-Provider LLM Routing Is Not a Problem, It's Your Architecture: Inference in Production Series

3 days ago

Why serious teams run multi-provider inference by default, and where DigitalOcean’s first-party...

What Kimi K3 Costs to Run

5 days ago

Introduction Kimi K3, released in July 2026 by Moonshot AI, is the...

What Inference Provider Lock-In Actually Looks Like in Production

5 days ago

Every LLM API vendor points to the same fact as proof that...

Long-context LLM serving: the real tradeoffs in memory, latency, cost, and accuracy

6 days ago

What “long context” means, and why supported is not the same as...

p50 vs p99 Latency: Why Median Benchmarks Mislead AI Agent Workloads

8 days ago

Most published inference benchmarks lead with one number. A median time to...

Migrating Your AI Cloud Inference Off Frontier Model Companies

12 days ago

Introduction When building an LLM (Large Language Model) application, you may have...

Best OpenAI-compatible inference APIs: drop-in alternatives for 2026

13 days ago

TL;DR: The leading OpenAI-compatible inference APIs in 2026 are, in alphabetical order...

Prompt Caching in Practice: From 7% to 74% Hit Rate(Inference in Production Series)

13 days ago

Prompt-caching mechanics plus a first-hand measurement on DigitalOcean Serverless (Anthropic-style explicit cache_control)...

The Hidden Cost of Output Token Pricing for Llama 3.3 70B

14 days ago

Introduction When comparing models for Llama 3.3 70B-class workloads, most cost analyses...