The exact pipeline behind every location result, when each step is allowed to answer, where the tool reliably fails, and the editorial rules behind our written guides.
2026/09/19
Most "AI location finder" sites give you an answer and no way to judge it. This page is the opposite: it explains what actually runs when you upload a photo, when each step is allowed to give an answer, and — more usefully — the specific situations where the result is likely to be wrong.
Read it before you trust any answer this site gives you, including the confident-looking ones.
When you submit an image, three checks decide whether there is anything to go on. They are tried in order of how trustworthy their answers are, and each one is allowed to answer only when it clears its own bar. If none does, the tool says it could not determine a location. It does not invent one.
The file is parsed for embedded EXIF data. If the camera wrote GPS coordinates into the file and nothing has stripped them, that is a direct measurement rather than a guess, so it takes precedence over everything else. The coordinates are sent to a map service to look up the name of the place.
The catch is how often it is absent. Facebook, Instagram, X, WhatsApp, and most messaging apps strip EXIF on upload, and a screenshot never had it in the first place. In practice, most photos people bring to a tool like this have no GPS tag — which is precisely why the other checks exist. You can inspect this yourself, without uploading anything, using our EXIF location viewer, which reads the file in your own browser.
The image is sent to Google Cloud Vision, which matches the scene against a large reference set of known places — the Eiffel Tower, Sydney Opera House, Machu Picchu, and tens of thousands of less famous but still catalogued structures. Its coordinates come from that reference set, not from an estimate. A match is only used when Vision's own score is at least 0.5; in our testing every correct match scored well above that, and the one wrong match below it.
This is the step responsible for the results that feel like magic. It is also the one with the sharpest cliff: a landmark five metres outside its reference set returns nothing at all, and an ordinary street almost always does.
If neither of the first two answers, the image goes to a vision-language model running on Cloudflare Workers AI. It is asked for a place name — never coordinates — along with the region and country, what the answer rests on, and how sure it is. Its answer is used only when all of these hold:
Text is where this step earns its place. Place names and signs are often the single most decisive clue in an otherwise anonymous photo, and they are exactly what a landmark recogniser cannot use. Results from this step are labelled on screen as read from the picture, and shown without a percentage, because the model's own sense of certainty is not a measured probability.
Whichever step answered, the place is resolved against map data for a canonical name, country and coordinates; a background summary is fetched from Wikipedia; and a static map is generated so you can see the answer in context.
Fifty photographs, one taken near an ordinary point in each of fifty cities on six continents —
Tartu, Kumasi, Arequipa, Pune, Townsville and so on — drawn from Wikimedia Commons, skipping
anything titled as a famous landmark. Every one records the coordinates it was taken at, so an
answer can be scored by how many kilometres it misses by. Before uploading we stripped the EXIF
data, so the tool could not simply read the GPS tag, and renamed each file to something neutral
like IMG_0001.jpg. Each photo went in through the same endpoint the box on our homepage uses.
We ran the whole set three times, because the model does not answer identically every time.
| Run | Answered | Within 25 km | 25–750 km out | More than 750 km out |
|---|---|---|---|---|
| 1 | 9 | 6 | 3 | 0 |
| 2 | 6 | 6 | 0 | 0 |
| 3 | 7 | 6 | 1 | 0 |
| Total of 150 chances | 22 | 18 | 4 | 0 |
So: on photographs like these the tool says nothing about six times in seven. When it does answer, it lands within 25 km about four times in five, and it did not once put an answer more than 750 km from where the photo was taken. The worst miss in the three runs was 132 km — a pedestrian street in Bendigo given as central Melbourne. Five photographs were answered correctly in all three runs; the rest of the answers came and went between runs.
Where the answers came from:
| Answers | Within 25 km | |
|---|---|---|
| Landmark database match | 5 | 5 |
| Vision model reading the picture | 17 | 13 |
| — of those, resting on legible text | 15 | 11 |
| — of those, resting on a named structure | 2 | 2 |
An earlier version of this page reported a single run on a different set of fifty photographs: 14 answered, 10 of them within 25 km. A second set, gathered the same way, did not reproduce it. With the model we were using then, that second set produced 16 answers of which five or six were within 25 km, and six answers landed more than 750 km out — a street in Penang given as Kagoshima, 4,286 km away; Tbilisi given as France; Chengdu given as Taiwan.
One sample of fifty was too small to publish a hit rate from, and we published one. That is the correction.
The failure was not in the checks around the model but in the model itself, so we compared the vision models available to us on the same photographs. The one now in use answered less often than its predecessor and was right more often, and stopped producing answers on the wrong continent. The table above is that model, measured after the change.
We would rather you know the failure modes than discover them at a bad moment.
| Situation | What happens | Why |
|---|---|---|
| Generic interiors — hotel rooms, offices, cafés | Usually no result, or a bad one | No landmark, no terrain, and text is often chain branding |
| Plain natural scenes — a beach, a forest, a field | Low confidence or wrong region | Thousands of coastlines look alike at this resolution |
| Night photos | Degraded across every step | Landmark matching and text reading both depend on visible detail |
| Close-ups and macro shots | No result | There is no scene to geolocate |
| Heavily edited, filtered, or upscaled images | Confidence drops, errors rise | Editing removes the fine detail the models key on |
| AI-generated images | Confidently wrong | The scene resembles a real place without being one |
| Very recent construction | Missed or misplaced | Reference sets and map data lag reality |
| Deliberately obscured locations | May still resolve | Treat this as a reason for caution, not a feature |
Two structural limitations are worth stating plainly. First, the confidence score is a model output, not a probability of being correct — a 0.9 does not mean nine times out of ten. Second, landmark coverage is uneven by geography: densely photographed parts of Europe, North America, and East Asia are far better represented than most of Africa, Central Asia, and rural South America, so expect worse results there.
Treat every answer as a hypothesis and spend two minutes testing it:
If steps 1 and 2 disagree with the answer, the answer is wrong. That happens, and we would rather you catch it than cite it.
The articles in our blog are a separate product from the tool, and they follow their own rules:
Uploaded images are used only to answer the question you asked. They are not used to train models, not published, and not sold. Original files are deleted after processing. The EXIF viewer runs entirely in your browser and never uploads your photo at all. The full detail, including our advertising and cookie disclosures, is in the Privacy Policy and Cookie Policy.
If you think a result is wrong, or you have found a failure mode that is not on this page, tell us at support@whereisthisplace.org. Reports of bad results are how this page gets more accurate.