<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>OpenStreetMap Archives - Urban Geo Analytics</title>
	<atom:link href="https://urbangeoanalytics.com/tag/openstreetmap/feed/" rel="self" type="application/rss+xml" />
	<link>https://urbangeoanalytics.com/tag/openstreetmap/</link>
	<description>Spatial Analysis, GeoAI &#38; Machine Learning</description>
	<lastBuildDate>Thu, 01 Oct 2026 06:21:30 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://urbangeoanalytics.com/wp-content/uploads/2025/11/cropped-logo-urban-geo_512-32x32.png</url>
	<title>OpenStreetMap Archives - Urban Geo Analytics</title>
	<link>https://urbangeoanalytics.com/tag/openstreetmap/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Do Open-Weight LLMs Reason From the Spatial Context They Are Given?</title>
		<link>https://urbangeoanalytics.com/llm-faithfulness-benchmark-geographic-reasoning/</link>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 06:19:39 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Urbanism]]></category>
		<category><![CDATA[benchmark]]></category>
		<category><![CDATA[faithfulness]]></category>
		<category><![CDATA[gemma]]></category>
		<category><![CDATA[GeoAI]]></category>
		<category><![CDATA[GHS-POP]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[Llama]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[network catchment]]></category>
		<category><![CDATA[open-weight models]]></category>
		<category><![CDATA[OpenStreetMap]]></category>
		<category><![CDATA[OSMnx]]></category>
		<category><![CDATA[Qwen3]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=3118</guid>

					<description><![CDATA[<p>A new paper asks a question geographic evaluation has mostly skipped: once a language model is handed the right local data, does it reason from it, or fall back on what it already believes about the place? Sixteen open-weight configurations, three cities, ten seeds, and one planted false premise per case.</p>
<p>The post <a href="https://urbangeoanalytics.com/llm-faithfulness-benchmark-geographic-reasoning/">Do Open-Weight LLMs Reason From the Spatial Context They Are Given?</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-1 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-0 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-1 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Language models know a lot about places, but they reason poorly over space and are unreliable when asked about a location from coordinates alone. The usual fix is to supply the relevant data in the prompt. This paper examines what happens next. Once the facts are in front of the model, does it use them, or does it override them with its own prior about the place?</p>
</div><div class="fusion-title title fusion-title-1 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">The paper</span></h2></div><div class="fusion-text fusion-text-2 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="19" data-line-type="context" data-line-index="18">
<div data-line="10" data-line-type="context" data-line-index="9">&#8220;Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning&#8221; separates two things that are usually measured together: being correct about the world, and being faithful to the context supplied.</div>
<div data-line="11" data-line-type="context" data-line-index="10"></div>
<div data-line="12" data-line-type="context" data-line-index="11">To do so, it fixes the context. For a point on a map, a short &#8220;spatial brief&#8221; is computed from OpenStreetMap and GHS-POP over the area reachable on foot: residents, density, buildings, street length, points of interest by category. The model receives finished numbers and only has to interpret them. Because the brief is known, every claim in an answer can be checked against it.</div>
<div data-line="13" data-line-type="context" data-line-index="12"></div>
<div data-line="14" data-line-type="context" data-line-index="13">The study applies this to three contrasting places, each tied to a theme in urban research that the models are likely to have met in training. The first is a neighbourhood on Chicago&#8217;s West Side, where 7,128 residents live within an 800 m walk that contains no supermarket or grocery store. The second is the historic core of Paris near Le Marais, with about 26,800 residents per km² and 971 points of interest within 400 m. The third is Hanoi&#8217;s Old Quarter, a district known for shops and tourism that nonetheless holds over 7,000 residents within a 300 m walk. The three are ordered by difficulty: Chicago asks the model to notice an absence, Paris to interpret a number, and Hanoi to hold a high resident count against a strong reputation.</div>
</div>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-1" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-1 hover-type-none"><img fetchpriority="high" decoding="async" width="1838" height="2000" title="fig1_pipeline" src="https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-scaled.png" alt class="img-responsive wp-image-3122" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-200x218.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-400x435.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-600x653.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-800x871.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-1200x1306.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-scaled.png 1838w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">The two-stages pipeline</div></div></div></div><div class="fusion-title title fusion-title-2 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-bottom:25px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">How faithfulness is measured</span></h2></div><div class="fusion-text fusion-text-3 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="27" data-line-type="context" data-line-index="26">Answers are split into atomic claims. Each claim is labelled by its source, the brief or the model&#8217;s own knowledge, and by its correctness. This gives four categories: brief-true, brief-false, recall and hallucination.</div>
<div data-line="28" data-line-type="context" data-line-index="27"></div>
<div data-line="29" data-line-type="context" data-line-index="28">Each conversation then ends with a trap. The user casually asserts something that sounds right but that the brief contradicts. In Chicago, that supermarkets are an easy walk away, in a catchment that has none. In Paris, that a district with 26,800 residents per km² is a calm, low-density corner. In Hanoi, that hardly anyone lives in the Old Quarter, where the brief counts over 7,000 residents within a 300 m walk.</div>
<div data-line="30" data-line-type="context" data-line-index="29"></div>
<div data-line="31" data-line-type="context" data-line-index="30">Sixteen configurations from Qwen3, Gemma 3, Gemma 4 and Llama 3.x, from 1B to 14B parameters, are each run ten times per case: 1,440 responses and about 22,000 labelled claims.</div>
</div><div class="fusion-title title fusion-title-3 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Main findings</span></h2></div><div class="fusion-text fusion-text-4 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="37" data-line-type="context" data-line-index="36">&#8211; Family matters more than size. Out of 30, Gemma 4 scores 22.5 to 23 on trap resistance and Qwen3 10 to 20.5. Gemma 3 scores 1 to 4 and Llama 2 to 6.5. The 1.7B Qwen3 outscores the 12B Gemma 3.</div>
<div data-line="38" data-line-type="context" data-line-index="37">&#8211; Reading and defending are different skills. Several models report every figure correctly, then drop them the moment the user disagrees.</div>
<div data-line="39" data-line-type="context" data-line-index="38">&#8211; The more plausible the premise, the weaker the resistance: 65% in Chicago, 43% in Paris, 14% in Hanoi.</div>
<div data-line="40" data-line-type="context" data-line-index="39">&#8211; Thinking mode makes answers 2.6 to 4.6 times slower without a consistent gain in grounding.</div>
<div data-line="41" data-line-type="context" data-line-index="40">&#8211; Outcomes change between identical runs, so a single generation is one draw, not a verdict.</div>
</div><iframe id="nscr-profiles"
  src="https://urbangeoanalytics.com/wp-content/uploads/2026/10/nscr-llm-model-profiles.html"
  title="Six-axis profile per model: grounding, hallucination, trap resistance, stability, concision and speed"
  style="width:100%;height:1040px;border:0;" loading="lazy"></iframe>
<script>
window.addEventListener('message', function (e) {
  if (e.data && e.data.nscrFigureHeight) {
    document.getElementById('nscr-profiles').style.height = (e.data.nscrFigureHeight + 10) + 'px';
  }
});
</script>
<p style="font-size:13px;text-align:center;">Six-axis profile per model. Use the buttons to switch between pooled and per-city profiles; hover a point for its value.</p><div class="fusion-title title fusion-title-4 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Why it matters</span></h2></div><div class="fusion-text fusion-text-5 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="37" data-line-type="context" data-line-index="36">
<div data-line="3" data-line-type="context" data-line-index="2">Scoring answers only against the truth would rate many of these models as reliable. For decision-support uses, the practical lesson is that a larger model that yields to its user is a worse choice than a smaller one that holds to the data.</div>
<div data-line="4" data-line-type="context" data-line-index="3"></div>
<div data-line="5" data-line-type="context" data-line-index="4">This matters most in geography, where what a model believes about a place is often a reputation rather than a fact: the Marais as a quiet historic quarter, Hanoi&#8217;s Old Quarter as shops and tourists rather than residents. Planners, analysts and public agencies are starting to put local data in front of these models precisely to get past such assumptions. If the model drops that data as soon as a confident user repeats the stereotype, the retrieval step has achieved nothing.</div>
<div data-line="6" data-line-type="context" data-line-index="5"></div>
<div data-line="7" data-line-type="context" data-line-index="6">The paper also shows that this risk cannot be read off a model card. Parameter count does not predict it, a reasoning mode does not fix it, and one good answer does not rule it out. It has to be measured, on the data and the questions the model will actually face.</div>
</div>
</div><div class="fusion-title title fusion-title-5 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:35px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">References and Links</span></h2></div><div class="fusion-text fusion-text-6 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><p>&#8211; Paper (preprint): <a class="keychainify-checked" href="https://arxiv.org/abs/2609.39437">https://arxiv.org/abs/2609.39437</a></p>
<p>&#8211; Code and data: <a class="keychainify-checked" href="https://github.com/perezjoan/NSCR-LLM">https://github.com/perezjoan/NSCR-LLM</a></p>
<p>&#8211; Archived release v1.0.0: <a class="keychainify-checked" href="https://doi.org/10.5281/zenodo.23056312">https://doi.org/10.5281/zenodo.23056312</a></p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-1 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-7"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--1" data-awb-toc-id="1" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-2 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/llm-faithfulness-benchmark-geographic-reasoning/">Do Open-Weight LLMs Reason From the Spatial Context They Are Given?</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
		
		
			</item>
	</channel>
</rss>
