Karakabakov
← All posts

Using Claude to crawl car prices across the Balkans

I wanted to know what a new car costs in Macedonia. All of them at once, so I could sort by price and see what a given budget buys. Then I wanted to know if the same car is cheaper in Serbia or Croatia.

Nobody publishes that. Each importer has its own website and its own price list, and the price list is usually a PDF. A few are tidy tables. Others are a brochure that opens with marketing pages and keeps the prices for the end, or a PowerPoint slide exported as a picture, so the file has no text in it at all. Then there are seven countries, five currencies, and languages in both Latin and Cyrillic.

So I built Avtoceni. It reads all of those documents once a week and turns them into one catalog at avtoceni.com. Claude does the reading. This post is about how. Most of what makes it work is ordinary code, and the model’s part is smaller than you might guess.

What does Claude actually do?

It gets the text of one document and returns the cars in it as JSON. That is all.

It does not browse. It does not decide which page to open next, it has no tools, and it never sees a URL it could follow. Fetching, finding links, caching, converting currencies and deciding what changed since last week are all plain TypeScript. I tried to keep the model’s job as small as I could, because a small job can be checked.

The rest of this post is the code around that one call.

Where do the pages come from?

From a file I keep by hand. Each country has a sources.yaml with one entry per importer:

- id: kia
  brand: Kia
  importer: Kia Motors Macedonia
  currency: EUR
  urls:
    - url: https://kia.mk/
      type: html
      note: homepage model grid with "од X €" prices
    - url: https://kia.mk/content/sportage-cenovnik-2027.pdf
    - url: https://kia.mk/content/ev3-cenovnik-2027.pdf

A URL can be marked follow: true, which means “find the PDF links on this page and read those too”, with a regex for which links count. That covers importers who rename the file every time they publish a new list.

There are 152 of these entries across the seven countries, 40 of them Macedonian. Which pages to read is a decision I would not hand to a crawler that wanders. The file also records what I looked at and left out. In Albania only two brands publish a current price list, so the Albanian file is mostly comments explaining why the others are missing.

How much of a page does the model see?

As little as possible. An importer’s homepage is mostly menus, scripts and cookie banners, so the HTML goes through a converter first. It drops scripts, styles, buttons, cookie banners and navigation, and it writes tables out row by row with | between the cells so a price grid survives as a grid.

There is a probe command that shows what would be sent without calling the model. This is Kia’s Macedonian homepage today:

html  final=https://www.kia.mk/  bytes=112683  text=1256 chars  price-like=10
  e.g. 17.500 € | 21.200 € | 23.950 € | 28.800 € | 29.950 € | 37.450 €
  4 document links (pdf)

112,683 bytes of HTML become 1,256 characters of text with ten prices in it. PDFs go through pdftotext -layout, which keeps the columns lined up. That matters more than it sounds: without the layout, a table with three price columns comes out as a list of numbers with no way to tell which is which.

What do you ask it for?

A schema. The call uses structured outputs with a zod schema, so the answer is either valid JSON in the shape I asked for or an error:

export const ExtractedVariantSchema = z.object({
  model: z.string(),
  variantName: z.string(),
  trim: z.string().nullable(),
  bodyType: z.enum(BODY_TYPES),
  fuel: z.enum(FUELS),
  transmission: z.enum(TRANSMISSIONS),
  powerKw: z.number().nullable(),
  price: z.number().nullable(),
  currency: z.enum(["EUR", "MKD", "RSD", "ALL", "BAM"]).nullable(),
  priceNote: z.string().nullable(),
  // ...
});

Almost every field can be null, and the prompt says so in its first rule: only include cars that appear in the text, never invent a variant, a price or a spec, and use null for anything the document does not state. A missing horsepower figure is fine. A made-up one is a bug I would never find.

Most of the prompt is about what a price is, because price lists disagree about it. A row often has a list price, a discount and a promotional price. The rule is that the price is the one you would pay today, and the other two go in a short note. Monthly instalments and leasing rates are never the price. Used cars, test cars and accessories are ignored.

Each country gets its own paragraph of vocabulary and traps. Bosnian lists print “KM” after a price, which is the convertible mark, and the prompt has to say it is not horsepower and not kilometres. Croatian lists build a price in four steps, and the last one includes a CO2 tax called PPMV. Serbian lists sometimes show a price without VAT next to one with it. None of that is clever. It is what a person learns from reading a pile of price lists, written down once.

What about PDFs that are pictures?

These need no separate OCR step.

The crawler counts the prices in the extracted text. If a PDF has fewer than three numbers with a currency next to them and fewer than ten numbers with thousands separators, the text layer is too thin to be a price list. It then renders up to twelve pages as JPEGs, 1,200 pixels wide, and sends those to the same prompt with the same schema.

It happens a lot. In the last Macedonian run, 52 of 243 PDFs needed page images. Across all seven countries it was 351 of 1,228.

How do you avoid paying for the same PDF every week?

Price lists change a few times a year, and the crawl runs every Monday. So every document is hashed, and the extracted cars are stored next to the hash. If the hash matches, the model is not called.

In the Macedonian run on 9 October, 258 of 374 URLs were unchanged.

Two details made the cache worth having. Some servers regenerate a PDF on each download, so the bytes differ while the content does not. For PDFs the cache keeps a hash of the extracted text as well, and either one matching counts. The other detail is that prices, technical figures and equipment lists are three separate calls, and the last two carry a version number in the cache. When I change the equipment prompt I raise its number, and the crawler reads every document again for equipment and leaves the prices alone.

Do you trust what comes back?

Not without checking it. The model’s answer goes through code that knows a few things about cars.

A price below 3,000 euros or above 500,000 is dropped, because it is an option pack or a typo. Technical figures have ranges too: a top speed outside 80 to 400 km/h, or a boot outside 50 to 3,000 litres, is thrown away. If the gearbox came back empty but the variant name says DSG or MT6, code fills it in.

Matching a car across borders has no model in it at all. Two variants are the same car only when the model, fuel and gearbox agree and the power is within 10%. Models that are spelled differently in two countries are paired in a file by hand, and a check command lists the near misses so I can add them.

When a price changes, the old one goes into a history file, and that is what the price changes page is built from.

What happens when a site says no?

Some importers’ firewalls answer 403 to GitHub’s servers and open fine from my desk. If that page was read successfully before, the crawl keeps what the last visit read, logs a warning and moves on. A page that returns 404 still fails, because that one is gone.

A separate health check fetches every URL the way the crawler would, without calling the model, and writes a report: which sources are failing, the likely reason and what to try. It runs after the crawl and fails its workflow when a source is broken, so GitHub emails me. The last report has one complaint: MINI’s price list times out.

What is the model, and where does it run?

The crawl is a GitHub Action on Monday at 03:00 UTC. It uses the Anthropic SDK with Claude Opus 5.5 at medium effort, with the system prompt cached, since it is identical for every document in a country. A smaller model, Haiku 4.5, has one side job: looking at candidate photos and saying whether each is a clean exterior shot of the car.

The output is JSON committed to the repository. The site is a static Astro build from that JSON, on Netlify, so a visitor never waits for a model and the site keeps working if a crawl fails.

On my own machine, with no API key set, the crawler falls back to the claude command line. That is handy for trying out a new source.

Not everything needed a model. MG’s site keeps its prices as JSON inside an HTML attribute on its configurator pages, so that importer has a 76-line parser of its own.

So are cars cheaper next door?

Mostly, yes, on paper. The regional page compares every model sold both in North Macedonia and in at least one neighbour. Today that is 167 models. 30 are cheapest at home and 137 are cheaper somewhere else, most often in Serbia.

The regional page on avtoceni.com: a table of models with the price in North Macedonia, the country where each is cheapest and the difference in percent

The biggest gap is a Volkswagen ID.4 at 58,764 euros in Macedonia against 40,118 in Montenegro. That is not money you can save by driving to Podgorica. VAT, excise and standard equipment differ between countries, and the page says so above the table. But it answers the question I started with.

The Macedonian price lists give 1,773 variants today, and all seven countries together about 9,600. The first commit is from 5 October.

What would I tell someone doing the same?

Keep the model away from the network, and give it one document and one schema at a time. Hash what you fetch so a quiet week costs almost nothing. Write the rules of your domain into the prompt and check the answers with code anyway. I would also build the command that shows you what the model would see before you build anything else. When an extraction is wrong, the text it was given is the first place to look.

Avtoceni is free at avtoceni.com, in Macedonian, English and Serbian. If a price looks wrong, every page has a button to report it.