<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Reducto Engineering</title><description>Technical writing from the team building Reducto</description><link>https://reducto.ai/</link><item><title>When does accepted speculative decoding translate into higher throughput and lower latency?</title><link>https://reducto.ai/engineering/posts/speculative-decoding-throughput-latency/</link><guid isPermaLink="true">https://reducto.ai/engineering/posts/speculative-decoding-throughput-latency/</guid><description>We benchmark multi-token prediction across workloads and concurrency levels to show when accepted speculative tokens improve throughput—and when verification overhead takes over.</description><pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate></item><item><title>The false positive problem: how LLM judges mislead table OCR benchmarks</title><link>https://reducto.ai/engineering/posts/llm-judges-table-ocr-false-positives/</link><guid isPermaLink="true">https://reducto.ai/engineering/posts/llm-judges-table-ocr-false-positives/</guid><description>We built a benchmark where ground truth is known by construction, then used it to show that frontier LLM judges hallucinate table OCR errors at non-trivial rates.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate></item><item><title>How We Cut Parsing Latency 3x by Building Our Own Orchestration System</title><link>https://reducto.ai/engineering/posts/streaq-scheduler-latency/</link><guid isPermaLink="true">https://reducto.ai/engineering/posts/streaq-scheduler-latency/</guid><description>Bursty, latency-sensitive traffic exposed the limits of per-tenant overflow pools. So we vendored a small task queue and built our own scheduling and traffic-management layer on top, keeping shared capacity warm, fair, and policy-aware while cutting parsing latency about 3x.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Fixing the Real Bottleneck in Our CV Serving Stack</title><link>https://reducto.ai/engineering/posts/fixing-cv-serving-bottleneck/</link><guid isPermaLink="true">https://reducto.ai/engineering/posts/fixing-cv-serving-bottleneck/</guid><description>Customer bursts exposed a weakness in our CV serving stack that average traffic mostly hid: the bottleneck was not model compute, but the pipeline around it.</description><pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate></item><item><title>How we built our agent-native observability stack</title><link>https://reducto.ai/engineering/posts/agent-native-observability/</link><guid isPermaLink="true">https://reducto.ai/engineering/posts/agent-native-observability/</guid><description>Most AI-for-observability demos start with a chatbot over logs. We cared about a less flashy question: could an agent understand our production system well enough to help operate it?</description><pubDate>Tue, 27 May 2025 00:00:00 GMT</pubDate></item></channel></rss>