AI & evaluation
Building the cross-border network.
Notes from RSA XB on cross-border logistics, AI tooling, and building software-first networks for freight forwarders, consolidators, and 3PLs.
AI & evaluationYour agent asked a human a question. The answer died with the task
A2A gives an agent a way to stop and ask. It says nothing about what happens to the answer afterwards. That gap is why AI triage pilots plateau — and it is fixable without a redeploy, if you are willing to let the agent publish what it can be taught.
Pavan Kumar TV15 min read
AI & evaluationIf the answer fits in a dropdown, it is not an agent problem
Smarter ops does not mean a model on every decision. Closed sets belong in registries. Judgment belongs where the schema runs out, and here is the test we use before anything calls a model.
Pavan Kumar TV7 min read
AI & evaluationLegal has specialist models. Logistics has one. That's the whole problem
Legal, healthcare and finance each have several specialist AI models with serious money behind them. We went looking for the logistics equivalent. We found one, and it's a research release.
Pavan Kumar TV5 min read
AI & evaluationYour shared mailbox isn't an inbox. It's an untyped API
Shared logistics mailboxes carry weight files, invoices, contracts, status reports and system noise in one stream. Classifying them is the easy half. Delivering them somewhere without redeploying is the hard half.
Parvez Alam6 min read
AI & evaluationAdding more context lowered classification accuracy
We assumed better retrieval would mean better classification. We tried dense embeddings, hybrid search, larger candidate sets. Accuracy went down. The bottleneck was somewhere else entirely.
Pavan Kumar TV5 min read
AI & evaluationThe benchmark we beat was 15.5% contaminated
31 of 200 test cases appear verbatim in the training split with identical labels. We found it while checking our own favourable result. Here's the de-leaked number.
Pavan Kumar TV5 min read
AI & evaluationFine-tuning scored 40%. Retrieval beat it without training
A fine-tuned 70B scores 40 percent on the published HTS benchmark. A training-free retrieval pipeline over the same public rulings scored 55. The interesting part isn't the gap. It's why the gap exists.
Pavan Kumar TV5 min read
AI & evaluationEveryone's accuracy number is unfalsifiable
Classification vendors advertise 90 to 96 percent. The only peer-reviewed number for the same task is 40. Those aren't contradictory findings. They're different measurements, and nobody says which one they ran.
Pavan Kumar TV5 min read
