HA Contact

Blog9 min read

5 hours a day to 30 minutes: building an AI tender radar for an engineering firm

How I turned hundreds of government tender notices a day into a short list an engineering team actually trusts, and why writing the code was the easy part.

A few months ago I started working with an engineering firm in India. They bid on government contracts for a living: power, solar, marine, civil and maintenance work. Every one of those contracts starts as a tender notice on a government portal, and there are dozens of portals.

Their team was finding tenders by hand. Someone would open each portal, search with the keywords they knew, download the documents and copy the good ones into a spreadsheet. They told me it took about 5 hours a day. And even then, shortlisted tenders were hard to keep track of, bids got missed because nothing sent a reminder, and there was no easy way to filter by buyer or by type of work.

Now they tell me it takes about 30 minutes. This is how I built it, and what went wrong along the way.

First, I listened

The first thing they gave me was a sheet of 12 sources to watch. Two of them returned nothing when I searched.

It took me a while to figure out why. The sheet used the word "portal" for three different things: the software a website runs on, the website itself, and the government body posting on it. One buyer posts under a completely different name, and another never uses its acronym. So before writing much code, I wrote a small vocabulary file (platform, portal, organisation), and we used those words the same way everywhere: in the code, in the database and with the client.

I also sent them 43 questions. One answer changed the whole design. I had assumed region should be a hard filter, but they actually sell nationwide. If I had built on my assumptions, the system would have hidden most of the work they want.

How I sketched the system

One line from our first conversation shaped everything: it has to run on its own, and nobody should have to open it to make it work. So before deciding on features, I sketched how it could fail.

System design of the tender radar: 25+ government portals feed seven stages. Collect hourly on AWS with Python; store and deduplicate in Supabase Postgres; AI screening through OpenRouter with Gemini and OpenAI; rank; read bid documents; price from the firm's own job history; track and notify in Slack with a dashboard on Vercel. Team decisions retrain the ranking every night. Sentry and Datadog watch every stage. Built with Claude Code, GitHub and GitHub Actions.
The system, stage by stage, with the tools at each step. Sentry and Datadog watch all of it.

I kept the shape simple on purpose: one pipeline, started every hour, that finishes in a few minutes. No queues, no microservices, nothing the firm doesn't need. On day one I built a "tracer bullet": one portal working end to end, from the listing to a report the team could read. Only after that did I add more portals and more stages.

To make it robust and autonomous:

  • Runs can be repeated safely. It remembers what it has already seen, so if a run crashes halfway, the next one just picks up. A lock makes sure two runs never overlap.
  • Every stage checks itself. If a portal suddenly extracts less than 80% of its fields, drops 15 points from its normal rate, or goes quiet for three hours, it gets flagged as broken instead of looking like a slow day.
  • Parts fail separately. The collector runs on a small server in India, the data sits in a managed database, and the dashboard is deployed on its own. If the dashboard breaks, collection keeps going.
  • Monitoring. I set up Sentry to catch errors and check in on every scheduled run, and Datadog for logs, metrics and monitors across the collector and the dashboard. If a run fails or a portal goes quiet, we get an alert in Slack before the team notices anything missing.
  • Backups that are tested. The database is backed up every night, encrypted, and I have actually restored from it to make sure it works.

Here is the journey of one tender through the system, using illustrative data from the public demo.

The life of one tender

Demo workspace, illustrative data

  1. Mon 09:12Published on the government e-marketplace: supply and commissioning of 11 kV indoor VCB panels.
  2. Mon 09:31Picked up by the hourly sweep and matched on “11 kV indoor VCB panels”, in the title and in item 1 of the bid document.
  3. Same runKept by the screen and ranked third of 412 open bids: a buyer they know, closing within a week.
  4. Tue 10:00Sent once in the morning digest, by email and in Slack, where it gets its own thread.
  5. TueShortlisted by the team. The bid detail already shows the deposit, eligibility and every line item.
  6. WedA suggested quote, computed from the firm’s own similar jobs, with the reasoning written out.
  7. Thu 15:00One day left: a reminder lands in the bid’s Slack thread.
  8. Fri 15:00The bid goes in before the deadline and moves to submitted.
  9. ResultWon or lost, the result is posted to the same thread, and the outcome feeds that night’s retraining.

Why only once an hour?

My first instinct was to check the portals as often as possible. But I measured first. Across 140 tenders, 60% were published around 6 pm, the median time from publishing to closing was 13 days, and only 1% closed within three days. So whether you see a tender at 6:05 pm or the next morning barely matters. What matters is seeing the right ones.

It's also about being polite. A full sweep is 76 requests, and doing that every second would mean millions of requests a day to government servers. The system makes about 444 requests a day, two seconds apart, and the worst-case delay on a new tender went from a day to an hour.

Fewer false positives, fewer false negatives

Any filter makes two kinds of mistakes. A false positive is a tender that shows up but doesn't fit: it wastes a minute of the team's time. A false negative is a tender that fits but never shows up: that can cost them a contract. These two are not equal, so I designed around the difference.

To reduce false negatives, I removed every place where a good tender could quietly disappear. Region became a ranking signal instead of a filter. Duplicates are joined on the tender reference, with title similarity only as a tie-breaker, because wrongly merging two tenders hides work while a duplicate only costs a glance. And the AI never deletes anything.

That last rule came from the team: people won't trust a filter they can't argue with. So every match saves its evidence: the exact text that matched, where it came from, and which version of the rules made the call. The AI screening can only set a tender aside when two models from different vendors agree, and each one has to quote the notice word for word. Anyone can bring it back with one click.

AI screening: one tender set aside with a quoted reason, another suggested with a quote
AI screening in the public demo workspace. Illustrative data.

To reduce false positives, I test every rule change against tenders the team has already judged. When I tightened the rules in September, my first attempt still let a USB cable through as "HT cable repair". The final version matched 5.42% of 20,000 tenders instead of 10.44% (about half the noise) and still kept all 33 of the 33 real tenders in the test set. A change only goes live if it improves one number without hurting the other.

Reading the bid documents

A tender title tells you very little. The real bid or no-bid decision is made on the bill of quantities, so the system downloads the documents and reads every line item. Not everything needs AI: on the biggest portal, the item list is just a small CSV file. What mattered more was measuring the boring parts. Switching PDF libraries took parsing from up to 48 seconds per document to under one, and testing on 658 real documents caught four kinds of extraction errors before the team ever saw them.

For shortlisted bids, it suggests a quote based on the firm's own past jobs, weighted by how similar they are, with past awards and competitors as extra context where available. The AI only writes the explanation; it never makes up the number. And if the inputs look wrong, it shows no number at all, because a wrong price is worse than none.

The bid detail view: closing date, why the tender matched, and the price of entry read from the bid document
Bid detail in the public demo workspace. Illustrative data, not the client's.

Ranking, and tracking every bid to the result

The killer feature for the team is the "recommended" order. Known buyers, invitations and close deadlines go to the top, and a model trained on their own bidding history fine-tunes the order. It scores an AUC of 0.71 on held-out data, so it's useful, not magic. Every night it retrains on the day's shortlists and rejections, including the reasons the team gives.

The tender list in recommended order, with known buyers, invitations and close deadlines first
The tender list in the public demo workspace. Illustrative data, not the client's.

Once a tender is shortlisted, the system tracks it through every stage: shortlisted, bidding, submitted, then won or lost. Each step sends an automated notification to their Slack channel, in one thread per tender, from the shortlist through the deadline reminders to the result. They can even react in Slack: a reaction records the decision, and removing it takes the decision back.

The morning digest: each new tender sent once, with the exact words it matched on
The morning digest, by email or Slack, in the public demo workspace. Illustrative data.

What broke

The hardest bugs were the ones where nothing looked broken:

  • One portal, when asked for page 5, quietly returned page 1. Now every page is checked by its content: page 5 has to start at row 41, or the run fails loudly.
  • One portal's timestamps end with "Z", which normally means UTC. They're not UTC.
  • Slack's API can reply 200 OK with "ok": false inside.
  • In September the dashboard list stopped updating, with zero errors. The cached list had grown to 2.72 MB, and the hosting platform silently refuses anything over 2 MB. Now the database decides what's fresh, and a separate watchdog compares the two.

Captchas were a different kind of wall. Some data, like past contract awards, sits behind them. I built a small OCR reader early on just to understand the problem, but it never runs in production. We never bypass a captcha. The system only reads where there isn't one, so the firm never has to worry about their own tool breaking a portal's rules.

The software engineering principles I followed

I used Claude Code a lot, with Matt Pocock's skills and gstack for planning, reviews and QA. But writing the code was never the hard part. These principles are what made it work:

  • Tracer bullet first: get one path working end to end before adding more.
  • One word, one meaning: a shared vocabulary for the code, the database and the client.
  • Write decisions down: 23 decision records so far, each with what we decided, why, and what we rejected.
  • Check the content, not the status code: a 200 response says nothing about the data inside it.
  • Silence is information: an unknown value is unknown, not zero, and an empty listing is something to investigate.
  • Tests that can actually fail: every change comes with a test that fails if the change is reverted, and a skipped test counts as a failure. The suite has over 2,300 tests now.
  • Safe database changes: add the new thing first, remove the old thing later, and make every migration safe to run twice.
  • Keep clients isolated: a second client now runs on the same platform, and a CI check makes sure nothing built for one touches the other.
  • Measure before optimising: every time I measured first, the answer was different from my guess.
  • A person decides: the software puts the evidence in front of the team. It never makes the bid decision for them.

The result

In the team's words, what used to take about 5 hours a day now takes about 30 minutes. They used to win around 10 bids a month. Last month they won 16, their best month so far. That's 60% more winning bids, and potentially 60% more revenue.

I take that last number with a pinch of salt, because wins also depend on their pricing, their people and the market. What I can say is that their time now goes into the bids worth winning, not into finding them.

If I did it again, I'd set up alerting on day one. For a short while, a failed overnight run told nobody, and for a system whose whole job is to not miss anything, a silent failure is the worst kind.

If you're building something similar, or you bid on government tenders and this 5-hours-a-day problem sounds familiar, I'd love to hear from you.