Reference
The error encyclopedia
Every entry answers three questions: what the error actually means, which of the possible causes you have, and the exact commands that fix it. Infrastructure and AI-stack errors sit side by side, because incidents stopped respecting that boundary.
Infrastructure
Kubernetes, Docker & cloud
imagepullbackoffImagePullBackOff in Kubernetes: What It Means and How to Fix Itexit-code-143Exit Code 143: What SIGTERM Means in Docker and Kubernetesexit-code-137Exit Code 137: SIGKILL, OOMKilled, and How to Fix BothoomkilledOOMKilled in Kubernetes: Why Pods Run Out of Memory and How to Fix ItcrashloopbackoffCrashLoopBackOff: How to Find Why Your Pod Keeps Restartingnode-not-readyKubernetes NodeNotReady: Diagnosis and Recoverykubernetes-exit-code-1Kubernetes Exit Code 1: Finding the Real Application Erroraws-lambda-timeoutAWS Lambda Timeout: Why "Task Timed Out" Happens and How to Fix It
AI stack
Models, GPUs & agent runtimes
openai-429-rate-limitOpenAI 429 Rate Limit Errors: What They Mean and How to Fix Themanthropic-529-overloadedAnthropic 529 overloaded_error: What It Means and How to Handle Itcontext-length-exceededcontext_length_exceeded: Why Your Prompt Overflows and How to Fix Itlitellm-errorsLiteLLM Proxy Errors: Diagnosing Auth, Config, and Routing Failurescuda-out-of-memoryCUDA Out of Memory (PyTorch): What It Means and How to Fix Itvllm-oomvLLM Out of Memory: Fixing Startup and Runtime OOMollama-errorsOllama Errors: Model Not Found, Port 11434, GPU and OOM Fixesmcp-server-errorsMCP Server Errors: Connection Failures, Causes and Fixes
Fixing the error is step one. Knowing which agent or teammate introduced it, and what it cost you, is Swfte Nexus.