The AI question
"Should we build this AI thing?" — the engagement where discovery most often ends in don't.
Representative stack
- Python
- PyTorch
- LangChain
- pgvector
- Postgres
- AWS
01 — The situation
Someone — a board, a competitor's press release, a demo that looked like magic — has raised the question, and now it needs a real answer: should this AI thing be built?
This is the engagement where our discovery most often ends in "don't build." That is the product working as designed. RAG, ML, and automation go into production only where they earn their cost; where they don't, you get the reasoning in writing and keep the rest of your budget.
02 — How the gates apply
Assay is the engagement, in miniature: we prototype against your real data, measure failure modes honestly, and put the build-or-don't recommendation in writing. The fee is the same either way, so the answer has no thumb on the scale.
If it's build: Smelt pins down evaluation criteria before the spec — what accuracy, latency, and cost make this worth running.
Temper is where AI systems earn trust: measured behavior on edge cases, in a report, before launch.
03 — What you receive
- WRITTEN RECOMMENDATION — BUILD, OR DON'T
- PROTOTYPE EVALUATED ON YOUR REAL DATA
- FAILURE-MODE ANALYSIS, MEASURED NOT VIBES
- RUNNING COST MODEL PER QUERY / PREDICTION
- IF BUILT: THE SYSTEM, ITS EVALS, AND ITS DOCS
04 — Where this goes wrong without discipline
- Starting from the technology instead of the task. "We need RAG" is not a problem statement.
- Demo-driven decisions. A demo hides its failure modes; production is made of them.
- No evaluation baseline. If nobody wrote down what good looks like, the model is always almost there.
Next — Start with discovery
Bring us the raw idea.
Discovery is a fixed fee. The recommendation is honest — even when it's “don't build.”

