Language models know a lot about places, but they reason poorly over space and are unreliable when asked about a location from coordinates alone. The usual fix is to supply the relevant data in the prompt. This paper examines what happens next. Once the facts are in front of the model, does it use them, or does it override them with its own prior about the place?

The paper

“Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning” separates two things that are usually measured together: being correct about the world, and being faithful to the context supplied.
To do so, it fixes the context. For a point on a map, a short “spatial brief” is computed from OpenStreetMap and GHS-POP over the area reachable on foot: residents, density, buildings, street length, points of interest by category. The model receives finished numbers and only has to interpret them. Because the brief is known, every claim in an answer can be checked against it.
The study applies this to three contrasting places, each tied to a theme in urban research that the models are likely to have met in training. The first is a neighbourhood on Chicago’s West Side, where 7,128 residents live within an 800 m walk that contains no supermarket or grocery store. The second is the historic core of Paris near Le Marais, with about 26,800 residents per km² and 971 points of interest within 400 m. The third is Hanoi’s Old Quarter, a district known for shops and tourism that nonetheless holds over 7,000 residents within a 300 m walk. The three are ordered by difficulty: Chicago asks the model to notice an absence, Paris to interpret a number, and Hanoi to hold a high resident count against a strong reputation.
The two-stages pipeline

How faithfulness is measured

Answers are split into atomic claims. Each claim is labelled by its source, the brief or the model’s own knowledge, and by its correctness. This gives four categories: brief-true, brief-false, recall and hallucination.
Each conversation then ends with a trap. The user casually asserts something that sounds right but that the brief contradicts. In Chicago, that supermarkets are an easy walk away, in a catchment that has none. In Paris, that a district with 26,800 residents per km² is a calm, low-density corner. In Hanoi, that hardly anyone lives in the Old Quarter, where the brief counts over 7,000 residents within a 300 m walk.
Sixteen configurations from Qwen3, Gemma 3, Gemma 4 and Llama 3.x, from 1B to 14B parameters, are each run ten times per case: 1,440 responses and about 22,000 labelled claims.

Main findings

– Family matters more than size. Out of 30, Gemma 4 scores 22.5 to 23 on trap resistance and Qwen3 10 to 20.5. Gemma 3 scores 1 to 4 and Llama 2 to 6.5. The 1.7B Qwen3 outscores the 12B Gemma 3.
– Reading and defending are different skills. Several models report every figure correctly, then drop them the moment the user disagrees.
– The more plausible the premise, the weaker the resistance: 65% in Chicago, 43% in Paris, 14% in Hanoi.
– Thinking mode makes answers 2.6 to 4.6 times slower without a consistent gain in grounding.
– Outcomes change between identical runs, so a single generation is one draw, not a verdict.

Six-axis profile per model. Use the buttons to switch between pooled and per-city profiles; hover a point for its value.

Why it matters

Scoring answers only against the truth would rate many of these models as reliable. For decision-support uses, the practical lesson is that a larger model that yields to its user is a worse choice than a smaller one that holds to the data.
This matters most in geography, where what a model believes about a place is often a reputation rather than a fact: the Marais as a quiet historic quarter, Hanoi’s Old Quarter as shops and tourists rather than residents. Planners, analysts and public agencies are starting to put local data in front of these models precisely to get past such assumptions. If the model drops that data as soon as a confident user repeats the stereotype, the retrieval step has achieved nothing.
The paper also shows that this risk cannot be read off a model card. Parameter count does not predict it, a reasoning mode does not fix it, and one good answer does not rule it out. It has to be measured, on the data and the questions the model will actually face.

References and Links

– Paper (preprint): https://arxiv.org/abs/2609.39437

– Code and data: https://github.com/perezjoan/NSCR-LLM

– Archived release v1.0.0: https://doi.org/10.5281/zenodo.23056312

Table of contents