On Où est le marché ?, a directory of open-air markets in France, anyone can add a market: its name, the town, the days and the opening hours. That's the whole idea of the site, and it works pretty well. But there's a catch: before a market goes live, the information has to be checked. A market listed on Sunday when it actually takes place on Tuesday means a visitor getting up early for nothing. And a site people stop trusting.
At first, I checked everything by hand. Find the town hall website, find the right page, read the opening hours, compare them with what was submitted... That's a good 5 minutes per market when everything goes well. With hundreds of markets waiting, the maths was quickly done: I needed a solution, fast, and above all cheap. This is a side project: paying for an API on every request was out of the question.
My solution: a small .NET console app that does the job for me, with two open-source tools running entirely on my laptop: SearXNG to search the web, and Ollama to run the AI.
TL;DR
- SearXNG (a self-hosted search engine) finds the pages that mention the market, Ollama (a local AI) extracts the days and opening hours from them.
- Everything runs locally, in Docker and on a laptop's graphics card: €0 in API costs, and no data sent to a third-party service.
- The AI must quote word for word the sentence where it found the hours, and plain old code checks that this sentence really is in the page. Otherwise, nothing gets published.
- A market is only published automatically when the source is reliable. Everything else goes into an Excel file for a manual check, which has become very quick.
Why not ChatGPT?
That's the first question people asked me. Three reasons:
- Cost. Each market means several web searches and several AI calls. Multiply that by hundreds of pending markets, then by 30,000 towns when you want to go further (more on that at the end): even at a few cents per call, the bill adds up fast. For a side project, that's a no.
- Web search. Google and Bing search APIs aren't free either. SearXNG is, and it runs at home.
- Control. I wanted to be able to run the job in the evening, and re-run it as many times as needed while fine-tuning, without counting. When it's local, you don't count.
My hardware: a laptop with an RTX 5060 (8 GB) graphics card. Nothing exceptional, but enough to run an 8-billion-parameter model.
The two tools
SearXNG is a "metasearch" engine: it queries several search engines (Google, Bing, DuckDuckGo, Wikipedia...) and returns the combined results as JSON. One command with Docker is all it takes to install it:
services:
searxng:
image: searxng/searxng:latest
container_name: searxng
restart: unless-stopped
ports:
- "8080:8080"
volumes:
- searxng-config:/etc/searxng
- searxng-cache:/var/cache/searxng
volumes:
searxng-config:
searxng-cache:
docker compose -f docker-compose.searxng.yml up -d
Ollama runs AI models (LLMs) on your own machine, behind a very simple HTTP API. I use qwen3:8b, a compact model that handles French very well:
ollama pull qwen3:8b
ollama serve
That's it for the setup. The rest is C#.
How it works, in 4 steps
For each market waiting to be validated:
- Search for the pages that mention it (town hall, tourist office, local press...).
- Keep only the useful passages of each page.
- Ask the AI to extract the days and opening hours.
- Check its answer with plain old code, before making a decision.
1. Searching with SearXNG
A simple HTTP request, with format=json:
var requestUri = $"search?q={Uri.EscapeDataString(query)}&format=json&language=fr-FR&safesearch=0";
var response = await client.GetAsync(requestUri, cancellationToken);
The query looks like marché Mauron Morbihan jours et horaires (the sources are French, so the queries are too). If the person who added the market gave a website, it's read first, without even going through the search.
2. Keeping the right passages
A town hall page contains plenty of opening hours: the town hall itself, the library, the swimming pool... Give the whole page to the AI and it may well pick the wrong ones (and it takes much longer to answer).
So before any call, I cut the page down to the passages around the word "marché" that contain a day or a time. A page with no such passage is discarded straight away, without calling the AI. It's simple, and it's what sped things up the most.
3. Asking the AI
The call to Ollama fits in a few lines:
var request = new
{
model = "qwen3:8b",
prompt,
stream = false,
format = "json", // JSON answer, ready to deserialize
think = false, // no "thinking out loud": faster
options = new
{
temperature = 0.1, // as little creativity as possible
num_ctx = 8192, // context size (see the pitfalls below)
},
};
var response = await client.PostAsJsonAsync("api/generate", request, cancellationToken);
And the prompt, in short:
Find in the excerpt below the weekly opening days and hours of the market in {town}.
Rules:
- Use only the excerpt. Never guess, never complete missing hours.
- Ignore other opening hours (town hall, shops, library...) and other towns.
- "evidence" must be copied word for word from the excerpt,
and contain every day and time you report.
Return this JSON only:
{
"scheduleFound": true,
"sessions": [ { "day": "mardi", "start": "08:00", "end": "12:30" } ],
"evidence": "verbatim quote",
"confidenceScore": 90
}
One important detail: I don't give the AI the hours entered by the user. In my first attempts, I gave them "to help"... and it simply confirmed them! Now it searches blind, and the code compares its result with the submission afterwards.
4. Checking (the most important part)
An LLM, even a well-guided one, can make things up. I've had a perfectly plausible opening time... that appeared nowhere on the page.
Hence the idea that makes the tool reliable: the AI must provide an exact quote (the evidence field), and plain old code, with no AI at all, checks that:
- the quote really exists in the page;
- every day reported is actually written in the quote (or covered by "du mardi au samedi", "tous les jours"...);
- every time reported is actually written in the quote.
if (!IsQuotedFrom(extraction.Evidence, pageText))
{
return null; // quote not found in the page: answer rejected
}
var quotedDays = FrenchCalendar.ExtractDays(extraction.Evidence);
if (extraction.Sessions.Any(session => !quotedDays.Contains(session.Day)))
{
return null; // a day that isn't in the quote: rejected
}
var quotedTimes = FrenchCalendar.ExtractTimes(extraction.Evidence);
if (extraction.Sessions.Any(s => !quotedTimes.Contains(s.Start!.Value) || !quotedTimes.Contains(s.End!.Value)))
{
return null; // a made-up time: rejected
}
When a check fails, the AI gets a second chance, with the exact reason for the rejection ("the time 14:00 is not in your quote"). But the checks themselves are never loosened.
Finally, each source gets a trust score: 100 for the town hall website, 95 for the tourist office, 80 for the press, 60 for a directory... A market is only published if its score reaches 85. In other words: the town hall website is enough on its own, a directory needs to be confirmed by another source.
The traps I fell into
Context size. By default, Ollama limits the context to 4,096 tokens, even when the model supports much more. A slightly long excerpt went over that limit: Ollama silently truncated the prompt, and the AI returned inconsistent JSON. The worst part? It mostly hit the richest pages. The fix: num_ctx = 8192 on every call.
Search engines don't like bots. After a few dozen automated queries, Google and co. block SearXNG. So I added a delay between requests and, for finding new markets (more on that below), sources that don't go through a search engine at all: the town hall's official website, found through the French administration directory (service-public.fr, free and keyless). For towns, SearXNG's Wikipedia engine is also far more tolerant.
A cache, or nothing. Every search, every page and every Ollama answer is stored in a small SQLite database. Re-running the job after a fix takes a few seconds instead of several hours.
A dry-run mode. With --dry-run, the tool does all the work but writes nothing to the database: it only produces the report. Essential for fine-tuning the rules safely.
The results
On my last batch of 200 markets waiting for validation:
- 6 markets were validated and published automatically, backed by an official source. And for 4 of them, the hours found were different from the ones submitted! For example, a market submitted on Sunday, while the tourist office says "every Tuesday from 8am to 12pm".
- The others weren't published automatically: source not reliable enough, hours not found, contradictory sources, several different time slots...
6 out of 200 may not sound like much. That's on purpose: I'd rather have a market left unpublished than a wrong one. And above all, the 194 others don't start from scratch. The tool generates an Excel report with, for each market, what it found, the exact quote, the source and its decision. The manual check no longer takes 5 minutes, just a few seconds: I read the quote, click the source if in doubt, and validate in the site's admin.
And since the tool knew how to look for a market... I also asked it to discover new ones. It went through nearly 30,000 towns that had no market at all on the site, across 88 départements, and added 662 of them, all confirmed by a reliable source: 640 thanks to the town hall's own website.
The cost of all this? €0, apart from the electricity of the laptop, which ran for a few evenings.
What I take away from it
- A local AI is more than enough to extract a precise piece of information from a text. No need for the biggest model out there: an 8-billion-parameter model on a consumer graphics card does the job very well.
- Never trust the AI blindly. Ask it to prove what it says (an exact quote) and check it with plain old code. That's what makes the difference between a demo and a tool you can run on thousands of rows.
- Prepare the ground. Giving it a short, targeted excerpt rather than a whole page improves both the speed and the quality of the answers.
- Keep a human in the loop, but only for the hard cases. The tool doesn't replace my check, it makes it ten times faster.
If you have a side project with a similar problem (data to check, sort or enrich), give Ollama a try: in an afternoon, you'll have an AI running at home, no credit card required.
And if you're ever looking for a market in France, you know where to go (the site is in French): ouestlemarche.fr 😉
A question about the code or the setup? Feel free to get in touch or leave a comment!

Comments (0)
No comments yet. Be the first to comment!
Leave a comment