Software Engineering & AI
Thorium Development Group brings 25+ years of software and systems engineering to the hard part of AI — the debugging, observability, and integration that turn a promising prototype into a system you can trust in production. We design, build, ship, and fix what breaks.
The Work, By the Numbers
No logos to show yet — something better: the actual bugs. Every one of these is a real (anonymized) production AI incident we found, and the fix it taught us.
An LLM pipeline hung for 36 minutes with no error and no log. Per-class timeouts and heartbeats now catch it in seconds.
Read the post →An emphatic “BE EXHAUSTIVE” made the model silently drop an entire section. Diagnosed by A/B prompt diffing and reversed.
A lossy “normalizer” model was corrupting inputs upstream. Feeding raw input downstream removed the errors at the source.
Moved judgment to author-time: the model writes a verified deterministic program once, then runs with zero per-transaction model calls.
Services
LLM integration, prompt engineering, and AI-assisted development workflows that actually work in production.
Governance frameworks, constraint enforcement, and audit trails for AI-assisted development.
ETL pipelines, vector databases, and data preparation for retrieval-augmented generation at scale.
Smart contract architecture, DeFi protocol design, and decentralized infrastructure engineering.
System design, scalable backends, and data-intensive application architecture.
Cloud architecture, CI/CD pipelines, containerization, and infrastructure automation.
Technical leadership, engineering strategy, and team scaling for companies that need senior guidance without a full-time hire.
Expertise
Distributed tracing for agent workflows, per-class timeouts, failing loud, and the instrumentation that makes non-deterministic AI systems debuggable in production.
25+ years designing backend systems that process large volumes of data. System design, scaling strategy, and the architectural decisions that compound.
Taming the LLM spend and latency blowups that surface at scale — metering baked into the client, right-sized models, and moving inference off the hot path where it doesn't belong.
Principal engineer, tech lead, and CTO experience across defense contracting and enterprise software. We build your team's capability, not a consulting dependency.
From the Blog
A production post-mortem on an LLM job that silently stalled for 36 minutes — why inherited SDK timeouts are dangerous, why AI calls come in latency classes, and the per-class timeout and heartbeat pattern that fixes it.
How to implement distributed tracing for multi-agent AI systems — propagating trace context across async boundaries, capturing LLM-specific signals, and building the observability that makes agent debugging possible.
A systematic approach to diagnosing tool call failures in AI agent systems — from incorrect parameter construction to silent schema mismatches and the debugging patterns that catch them.
A 30-minute strategy call to discuss your current technical needs and whether an engagement makes sense. No pitch deck. No sales pressure.