Hardware startup — RAG over internal runbooks
MQTT field data → Cloud Storage → chunked runbook corpus → pgvector → OpenAI function-calling agent. 62% of support tickets auto-resolved.
Custom RAG pipelines, LLM evals, voice agents, Copilot extensions, and MLOps that doesn’t break when a model version changes. We build the boring parts — retries, fallbacks, evals, observability — that turn a demo into a product.
Rate $80 – $100 USD / hour · fixed-fee and retainer available
Discovery (2 weeks, fixed fee)
We read your data, your existing prompts, and your failure modes. We write a one-page “what good looks like” and a cost model.
Build (4–8 weeks)
RAG pipeline, eval harness, and the first production agent. We ship behind a feature flag.
Operate (ongoing)
Optional retainer: weekly model-eval runs, monthly cost/perf review, on-call for production incidents.
LLM providers
OpenAI, Anthropic, Gemini, Mistral, vLLM, Ollama, Vertex AI
Frameworks
LangGraph, LlamaIndex, DSPy, Haystack, custom orchestration
Vector stores
Pinecone, Weaviate, pgvector, Qdrant, Vertex AI Vector Search
Evals
Promptfoo, Braintrust, custom labeled sets, golden traces
Voice
Twilio + Deepgram + Cartesia, Vapi, custom pipelines
Copilot / M365
Declarative agents, Teams message extensions, Graph connectors
Anonymized patterns from real engagements. Client names omitted; details available under NDA.
MQTT field data → Cloud Storage → chunked runbook corpus → pgvector → OpenAI function-calling agent. 62% of support tickets auto-resolved.
Declarative agent + Graph connector over a private DMS, deployed to a regulated tenant. Passed the Microsoft security review on first submission.
Bilingual (English + Hindi) voice agent handling 1,200 calls/day. p95 latency < 1.4s. Cost per call < $0.12.
Evals against a frozen golden set, model-version pinning with explicit upgrade windows, output-schema validation, and observability on cost and latency per call.
OpenAI, Anthropic, Google Gemini, Mistral, open-source via vLLM/Ollama, and on-prem via Vertex AI. We are not a reseller — we recommend per workload.
Yes. Declarative agents, message-extension plugins, and Graph connectors for the Microsoft 365 ecosystem, including the security review process.
Document audit → chunking + embedding strategy → vector store → retrieval + reranking → evals on a labeled set → deploy with usage telemetry.
When it actually helps. Most startups do not need fine-tuning — good retrieval, good prompts, and a long enough context window get you 90% of the way. We tell you honestly when fine-tuning is the right call.
A 30-minute call. No deck, no pitch — we read your repo or your architecture diagram and tell you what’s realistic.