<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Urban Geo Analytics</title>
	<atom:link href="https://urbangeoanalytics.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://urbangeoanalytics.com/</link>
	<description>Spatial Analysis, GeoAI &#38; Machine Learning</description>
	<lastBuildDate>Fri, 02 Oct 2026 06:50:31 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://urbangeoanalytics.com/wp-content/uploads/2025/11/cropped-logo-urban-geo_512-32x32.png</url>
	<title>Urban Geo Analytics</title>
	<link>https://urbangeoanalytics.com/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>Do Open-Weight LLMs Reason From the Spatial Context They Are Given?</title>
		<link>https://urbangeoanalytics.com/llm-faithfulness-benchmark-geographic-reasoning/</link>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Thu, 01 Oct 2026 06:19:39 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Urbanism]]></category>
		<category><![CDATA[benchmark]]></category>
		<category><![CDATA[faithfulness]]></category>
		<category><![CDATA[gemma]]></category>
		<category><![CDATA[GeoAI]]></category>
		<category><![CDATA[GHS-POP]]></category>
		<category><![CDATA[hallucination]]></category>
		<category><![CDATA[Llama]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[network catchment]]></category>
		<category><![CDATA[open-weight models]]></category>
		<category><![CDATA[OpenStreetMap]]></category>
		<category><![CDATA[OSMnx]]></category>
		<category><![CDATA[Qwen3]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=3118</guid>

					<description><![CDATA[<p>A new paper asks a question geographic evaluation has mostly skipped: once a language model is handed the right local data, does it reason from it, or fall back on what it already believes about the place? Sixteen open-weight configurations, three cities, ten seeds, and one planted false premise per case.</p>
<p>The post <a href="https://urbangeoanalytics.com/llm-faithfulness-benchmark-geographic-reasoning/">Do Open-Weight LLMs Reason From the Spatial Context They Are Given?</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-1 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-0 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-1 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Language models know a lot about places, but they reason poorly over space and are unreliable when asked about a location from coordinates alone. The usual fix is to supply the relevant data in the prompt. This paper examines what happens next. Once the facts are in front of the model, does it use them, or does it override them with its own prior about the place?</p>
</div><div class="fusion-title title fusion-title-1 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">The paper</span></h2></div><div class="fusion-text fusion-text-2 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="19" data-line-type="context" data-line-index="18">
<div data-line="10" data-line-type="context" data-line-index="9">&#8220;Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning&#8221; separates two things that are usually measured together: being correct about the world, and being faithful to the context supplied.</div>
<div data-line="11" data-line-type="context" data-line-index="10"></div>
<div data-line="12" data-line-type="context" data-line-index="11">To do so, it fixes the context. For a point on a map, a short &#8220;spatial brief&#8221; is computed from OpenStreetMap and GHS-POP over the area reachable on foot: residents, density, buildings, street length, points of interest by category. The model receives finished numbers and only has to interpret them. Because the brief is known, every claim in an answer can be checked against it.</div>
<div data-line="13" data-line-type="context" data-line-index="12"></div>
<div data-line="14" data-line-type="context" data-line-index="13">The study applies this to three contrasting places, each tied to a theme in urban research that the models are likely to have met in training. The first is a neighbourhood on Chicago&#8217;s West Side, where 7,128 residents live within an 800 m walk that contains no supermarket or grocery store. The second is the historic core of Paris near Le Marais, with about 26,800 residents per km² and 971 points of interest within 400 m. The third is Hanoi&#8217;s Old Quarter, a district known for shops and tourism that nonetheless holds over 7,000 residents within a 300 m walk. The three are ordered by difficulty: Chicago asks the model to notice an absence, Paris to interpret a number, and Hanoi to hold a high resident count against a strong reputation.</div>
</div>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-1" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-1 hover-type-none"><img fetchpriority="high" decoding="async" width="1838" height="2000" title="fig1_pipeline" src="https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-scaled.png" alt class="img-responsive wp-image-3122" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-200x218.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-400x435.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-600x653.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-800x871.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-1200x1306.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/10/fig1_pipeline-scaled.png 1838w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">The two-stages pipeline</div></div></div></div><div class="fusion-title title fusion-title-2 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-bottom:25px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">How faithfulness is measured</span></h2></div><div class="fusion-text fusion-text-3 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="27" data-line-type="context" data-line-index="26">Answers are split into atomic claims. Each claim is labelled by its source, the brief or the model&#8217;s own knowledge, and by its correctness. This gives four categories: brief-true, brief-false, recall and hallucination.</div>
<div data-line="28" data-line-type="context" data-line-index="27"></div>
<div data-line="29" data-line-type="context" data-line-index="28">Each conversation then ends with a trap. The user casually asserts something that sounds right but that the brief contradicts. In Chicago, that supermarkets are an easy walk away, in a catchment that has none. In Paris, that a district with 26,800 residents per km² is a calm, low-density corner. In Hanoi, that hardly anyone lives in the Old Quarter, where the brief counts over 7,000 residents within a 300 m walk.</div>
<div data-line="30" data-line-type="context" data-line-index="29"></div>
<div data-line="31" data-line-type="context" data-line-index="30">Sixteen configurations from Qwen3, Gemma 3, Gemma 4 and Llama 3.x, from 1B to 14B parameters, are each run ten times per case: 1,440 responses and about 22,000 labelled claims.</div>
</div><div class="fusion-title title fusion-title-3 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Main findings</span></h2></div><div class="fusion-text fusion-text-4 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="37" data-line-type="context" data-line-index="36">&#8211; Family matters more than size. Out of 30, Gemma 4 scores 22.5 to 23 on trap resistance and Qwen3 10 to 20.5. Gemma 3 scores 1 to 4 and Llama 2 to 6.5. The 1.7B Qwen3 outscores the 12B Gemma 3.</div>
<div data-line="38" data-line-type="context" data-line-index="37">&#8211; Reading and defending are different skills. Several models report every figure correctly, then drop them the moment the user disagrees.</div>
<div data-line="39" data-line-type="context" data-line-index="38">&#8211; The more plausible the premise, the weaker the resistance: 65% in Chicago, 43% in Paris, 14% in Hanoi.</div>
<div data-line="40" data-line-type="context" data-line-index="39">&#8211; Thinking mode makes answers 2.6 to 4.6 times slower without a consistent gain in grounding.</div>
<div data-line="41" data-line-type="context" data-line-index="40">&#8211; Outcomes change between identical runs, so a single generation is one draw, not a verdict.</div>
</div><iframe id="nscr-profiles"
  src="https://urbangeoanalytics.com/wp-content/uploads/2026/10/nscr-llm-model-profiles.html"
  title="Six-axis profile per model: grounding, hallucination, trap resistance, stability, concision and speed"
  style="width:100%;height:1040px;border:0;" loading="lazy"></iframe>
<script>
window.addEventListener('message', function (e) {
  if (e.data && e.data.nscrFigureHeight) {
    document.getElementById('nscr-profiles').style.height = (e.data.nscrFigureHeight + 10) + 'px';
  }
});
</script>
<p style="font-size:13px;text-align:center;">Six-axis profile per model. Use the buttons to switch between pooled and per-city profiles; hover a point for its value.</p><div class="fusion-title title fusion-title-4 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Why it matters</span></h2></div><div class="fusion-text fusion-text-5 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><div data-line="37" data-line-type="context" data-line-index="36">
<div data-line="3" data-line-type="context" data-line-index="2">Scoring answers only against the truth would rate many of these models as reliable. For decision-support uses, the practical lesson is that a larger model that yields to its user is a worse choice than a smaller one that holds to the data.</div>
<div data-line="4" data-line-type="context" data-line-index="3"></div>
<div data-line="5" data-line-type="context" data-line-index="4">This matters most in geography, where what a model believes about a place is often a reputation rather than a fact: the Marais as a quiet historic quarter, Hanoi&#8217;s Old Quarter as shops and tourists rather than residents. Planners, analysts and public agencies are starting to put local data in front of these models precisely to get past such assumptions. If the model drops that data as soon as a confident user repeats the stereotype, the retrieval step has achieved nothing.</div>
<div data-line="6" data-line-type="context" data-line-index="5"></div>
<div data-line="7" data-line-type="context" data-line-index="6">The paper also shows that this risk cannot be read off a model card. Parameter count does not predict it, a reasoning mode does not fix it, and one good answer does not rule it out. It has to be measured, on the data and the questions the model will actually face.</div>
</div>
</div><div class="fusion-title title fusion-title-5 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:35px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">References and Links</span></h2></div><div class="fusion-text fusion-text-6 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><p>&#8211; Paper (preprint): <a class="keychainify-checked" href="https://arxiv.org/abs/2609.39437">https://arxiv.org/abs/2609.39437</a></p>
<p>&#8211; Code and data: <a class="keychainify-checked" href="https://github.com/perezjoan/NSCR-LLM">https://github.com/perezjoan/NSCR-LLM</a></p>
<p>&#8211; Archived release v1.0.0: <a class="keychainify-checked" href="https://doi.org/10.5281/zenodo.23056312">https://doi.org/10.5281/zenodo.23056312</a></p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-1 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-7"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--1" data-awb-toc-id="1" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-2 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/llm-faithfulness-benchmark-geographic-reasoning/">Do Open-Weight LLMs Reason From the Spatial Context They Are Given?</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
		
		
			</item>
		<item>
		<title>UVLM v4.0.0 — Gemma 4, the Transformers 5 Migration, and Why This One Is a Major Version</title>
		<link>https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/</link>
					<comments>https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Mon, 17 Aug 2026 07:45:10 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Package]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[gemma]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=3029</guid>

					<description><![CDATA[<p>Highlights  New model family: Gemma 4 (Google DeepMind, released April 2026) joins as the fifth family — E2B, E4B, and 12B Instruct, bringing the registry to 24 checkpoints Breaking change, done honestly: UVLM now requires Transformers ≥ 5.15; all four existing families were re-validated on GPU before release, and v3.2.0 remains installable  [...]</p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/">UVLM v4.0.0 — Gemma 4, the Transformers 5 Migration, and Why This One Is a Major Version</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-2 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-2 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-3 hover-type-none"><img decoding="async" width="1536" height="1024" title="UVLM 4.0.0" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0.png" alt class="img-responsive wp-image-3043" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0.png 1536w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-3 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-8 fusion-text-no-margin" style="--awb-margin-bottom:-10px;"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-9" style="--awb-margin-top:-20px;"><ul>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>New model family:</strong> Gemma 4 (Google DeepMind, released April 2026) joins as the fifth family — E2B, E4B, and 12B Instruct, bringing the registry to 24 checkpoints</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Breaking change, done honestly:</strong> UVLM now requires Transformers ≥ 5.15; all four existing families were re-validated on GPU before release, and v3.2.0 remains installable for Transformers 4.x environments</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Measured, not guessed:</strong> Gemma 4&#8217;s &#8220;effective parameters&#8221; hide ~10–16 GB raw checkpoints — this release documents exactly what runs on an 8 GB GPU, and how</li>
</ul>
</div><div class="fusion-title title fusion-title-6 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Why a major version?</span></h2></div><div class="fusion-text fusion-text-10 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><a class="keychainify-checked" href="https://github.com/perezjoan/UVLM">UVLM</a> has followed one rule since the package release: a new model family is a minor version, because it breaks nothing. v4.0.0 breaks that streak for a reason we could not code around: <strong>Gemma 4 does not exist in any Transformers 4.x release.</strong> We verified this empirically — 4.57.6 is the final version of the 4.x line, and it does not register the <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">gemma4</code> architecture; support begins in the 5.x line. Adopting the family therefore means lifting the <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">transformers &lt; 5.0.0</code> cap that UVLM has carried since v3.0.1.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">That cap was not decoration. It existed because early Transformers 5.x releases crashed Qwen2.5-VL at load time with a weight-conversion error. So before this release, all four existing families — LLaVA-NeXT, Qwen2.5-VL, Qwen3-VL, InternVL3.5 — were re-validated on GPU under Transformers 5.15, in 4-bit, on real inference tasks. The historical Qwen2.5-VL crash is <strong>confirmed fixed</strong>: the 7B model loads and answers correctly at full speed. That validation is what makes this a release rather than a gamble.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">If your environment must stay on Transformers 4.x, nothing is taken from you: <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">pip install git+https://github.com/perezjoan/UVLM.git@v3.2.0</code> pins the last 4.x-compatible release, permanently.</p>
</div><div class="fusion-title title fusion-title-7 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What is Gemma 4?</span></h2></div><div class="fusion-text fusion-text-11 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Gemma 4 is Google DeepMind&#8217;s latest open multimodal generation, released in April 2026 under Apache 2.0. UVLM v4.0.0 integrates three Instruct checkpoints:</p>
</div>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Model</th>
<th align="left"> Parameters</th>
<th align="left"> Raw checkpoint</th>
<th align="left"> Runs on</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Gemma 4 E2B Instruct</td>
<td align="left"> ~2B effective</td>
<td align="left"> ~10 GB</td>
<td align="left"> 8 GB GPU in FP16 with CPU offload</td>
</tr>
<tr>
<td align="left">Gemma 4 E4B Instruct</td>
<td align="left"> ~4B effective</td>
<td align="left"> ~16 GB</td>
<td align="left">Larger-VRAM environments (Colab A100/L4)</td>
</tr>
<tr>
<td align="left">Gemma 4 12B Instruct</td>
<td align="left"> 12B</td>
<td align="left"> ~24 GB</td>
<td align="left">Larger-VRAM environments (Colab A100/L4)</td>
</tr>
</tbody>
</table>
</div>
<div class="fusion-text fusion-text-12 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Notice the third column, because it is this release&#8217;s most useful finding. E2B and E4B are <strong>&#8220;effective&#8221;-parameter models</strong>: Per-Layer Embeddings give them the <em>compute</em> profile of a 2B/4B model, but the embedding tables push the <em>raw</em> checkpoint far beyond what the name suggests. A &#8220;2B&#8221; model that downloads 10 GB of weights behaves very differently from Qwen3-VL 2B&#8217;s genuinely small footprint — and honest benchmarking infrastructure should say so, with numbers.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Technically, Gemma 4 shares the tokenizing-chat-template pipeline introduced with InternVL3.5, with one addition: Gemma 4 models other than E2B/E4B wrap their output in thought-channel tags even when thinking is disabled, and the backend strips them automatically. As always, the family appeared in the notebook selector with zero interface changes.</p>
</div><div class="fusion-title title fusion-title-8 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What actually runs on a laptop GPU</span></h2></div><div class="fusion-text fusion-text-13 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>We validated Gemma 4 on an 8 GB RTX 5060, and the result inverts the usual intuition: <strong>FP16 is the low-memory mode.</strong> In FP16, the compute-heavy layers stay on the GPU while the PLE embedding tables — lookup-only structures designed to live off-accelerator — offload to system RAM in half precision. Measured: about 17 seconds per image. Slow, but fully functional.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-4 hover-type-none"><img decoding="async" width="1215" height="641" title="gemma illustration" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration.png" alt class="img-responsive wp-image-3039" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-200x106.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-400x211.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-600x317.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-800x422.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-1200x633.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration.png 1215w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div><div class="fusion-text fusion-text-14 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>4-bit quantization — normally the memory-saver — fails here, for a subtle reason: offloaded modules are kept in FP32, roughly doubling the RAM requirement, and any further spill to disk is unsupported by bitsandbytes. Rather than leave users with a 200-line traceback, v4.0.0 detects this case and raises a two-sentence error recommending FP16. To support all of this, the loader gained general CPU-offload capability for oversized checkpoints — a change that only <em>permits</em> offload: models that fit entirely on the GPU are placed exactly as before.</p>
</div><div class="fusion-title title fusion-title-9 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Also in this release</span></h2></div><div class="fusion-text fusion-text-15 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>A failed model load in the notebooks now invalidates the previously loaded model, so a batch run after a failed load errors out loudly instead of silently benchmarking the wrong checkpoint — a trap we fell into ourselves during validation, and one that the per-model output filenames from v3.2.0 caught. InternVL3.5 users on Transformers 5 will see a harmless &#8220;tied weights&#8221; warning caused by an upstream config inconsistency; Transformers resolves it correctly.</p>
</div><div class="fusion-title title fusion-title-10 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Getting started</span></h2></div><div class="fusion-text fusion-text-16" style="--awb-margin-top:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="Bash" data-enlighter-title="Bash">pip install --upgrade --force-reinstall git+https://github.com/perezjoan/UVLM.git</pre>
</div><div class="fusion-text fusion-text-17 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><p>Note: no <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">--no-deps</code> this time — the whole point of the upgrade is that pip pulls Transformers 5.15 for you. Colab users get v4.0.0 automatically on their next session, since the notebook always installs the latest version; the re-validation above is what makes that automatic jump safe. The three-block workflow, consensus validation, chain-of-thought mode, and truncation detection all work with Gemma 4 out of the box.</p>
</div><div class="fusion-title title fusion-title-11 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Where UVLM stands</span></h2></div><div class="fusion-text fusion-text-18 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Five months ago, UVLM was a two-family package. It now supports <strong>five families and 24 checkpoints from 1B to 110B parameters</strong> — LLaVA-NeXT, Qwen2.5-VL, Qwen3-VL, InternVL3.5, Gemma 4 — behind one interface, one prompt format, one evaluation protocol. Three families were added in three releases without a single notebook edit, each validated on hardware before shipping. That is the registry we will be benchmarking against in upcoming applied work — more on that soon.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Full change log in <a class="keychainify-checked" href="https://github.com/perezjoan/UVLM/blob/main/VERSIONS.txt">VERSIONS.txt</a> · Source and releases on GitHub · If you use UVLM in research, please cite our Software paper.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-4 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-19"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--2" data-awb-toc-id="2" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-5 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-3 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"></div></div></p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/">UVLM v4.0.0 — Gemma 4, the Transformers 5 Migration, and Why This One Is a Major Version</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>UVLM v3.2.0 — InternVL3.5 Joins the Registry, With Zero Notebook Changes</title>
		<link>https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/</link>
					<comments>https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Wed, 12 Aug 2026 08:14:01 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Package]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[InternVL]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2982</guid>

					<description><![CDATA[<p>UVLM v3.2.0 adds InternVL3.5 (1B–38B, six checkpoints): 21 open VLM checkpoints across 4 families, one Python interface. The new family appeared in the notebooks without a single notebook edit — plus per-model output files for cleaner benchmarking.</p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/">UVLM v3.2.0 — InternVL3.5 Joins the Registry, With Zero Notebook Changes</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-4 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-5 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-6 hover-type-none"><img decoding="async" width="1536" height="1024" title="uvlm3.2.0 illustration" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration.png" alt class="img-responsive wp-image-2991" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration.png 1536w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-6 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-title title fusion-title-12 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Highlights</span></h2></div><div class="fusion-text fusion-text-20"><ul>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>New model family:</strong> InternVL3.5 (OpenGVLab, released August 2025) joins LLaVA-NeXT, Qwen2.5-VL, and Qwen3-VL — six checkpoints from 1B to 38B</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Zero notebook changes:</strong> the new family appeared in the selector automatically — the extensibility promise from v3.1.0, kept</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Per-model output files:</strong> each checkpoint now writes its own CSV, so resume mode can never mix results from different models</li>
</ul>
</div><div class="fusion-title title fusion-title-13 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">1. What is InternVL3.5?</span></h2></div><div class="fusion-text fusion-text-21 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">In <a class="keychainify-checked" href="https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/">v3.1.0</a> we added Qwen3-VL and made a promise: thanks to the new <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">FAMILY_GROUPS</code> registry, future model families would appear in the notebooks automatically, with no interface edits at all. Version 3.2.0 is that promise kept. <strong>InternVL3.5</strong> — the latest generation of OpenGVLab&#8217;s InternVL line, released in August 2025 — is now the fourth family in the registry, and neither notebook changed by a single line to display it.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">UVLM integrates the six Transformers-native <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">-HF</code> checkpoints, which run through the standard Transformers stack without any custom remote code. None of them is gated: no Hugging Face token required.</p>
</div>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Parameters</th>
<th align="left">VRAM (4-bit)</th>
<th align="left"> Typical hardware</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">InternVL3.5 1B</td>
<td align="left"> 1B</td>
<td align="left">~1 GB</td>
<td align="left">Any modern laptop GPU, free Colab T4</td>
</tr>
<tr>
<td align="left">InternVL3.5 2B</td>
<td align="left">2B</td>
<td align="left">~2 GB</td>
<td align="left">Any modern laptop GPU, free Colab T4</td>
</tr>
<tr>
<td align="left">InternVL3.5 4B</td>
<td align="left">4B</td>
<td align="left">~3 GB</td>
<td align="left">T4, RTX 3060</td>
</tr>
<tr>
<td align="left">InternVL3.5 8B</td>
<td align="left">8B</td>
<td align="left">~6 GB</td>
<td align="left">T4, RTX 4060/5060</td>
</tr>
<tr>
<td align="left">InternVL3.5 14B</td>
<td align="left">14B</td>
<td align="left">~9 GB</td>
<td align="left">L4, RTX 4070</td>
</tr>
<tr>
<td align="left">InternVL3.5 38B</td>
<td align="left">38B</td>
<td align="left">~22 GB</td>
<td align="left">A100, RTX 4090</td>
</tr>
</tbody>
</table>
</div>
<div class="fusion-text fusion-text-22 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The registry now totals <strong>21 checkpoints across 4 families</strong>, from 1B to 110B parameters — and the 1B entry replaces Qwen3-VL 2B as the smallest model UVLM has ever supported.</p>
</div><div class="fusion-title title fusion-title-14 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">2. A genuinely different pipeline</span></h2></div><div class="fusion-text fusion-text-23 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">InternVL3.5 is not a variation on the Qwen conventions — it uses the standard Transformers pattern in which the <strong>chat template tokenizes directly</strong> (<code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">apply_chat_template(tokenize=True)</code>), the generated tokens are sliced off after the prompt, and only the generated portion is decoded. That makes it the third distinct inference path in UVLM, alongside LLaVA&#8217;s string-based cleaning and Qwen&#8217;s separate vision preprocessing with token trimming. As always, all three converge at the same unified response parser — from the user&#8217;s side, InternVL3.5 is simply one more family in the dropdown.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-7 hover-type-none"><img decoding="async" width="1308" height="644" title="internvl uvlm" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm.png" alt class="img-responsive wp-image-2983" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-200x98.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-400x197.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-600x295.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-800x394.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-1200x591.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm.png 1308w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div><div class="fusion-text fusion-text-24 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Loading follows the same BF16-aware logic introduced in v3.1.0: BF16 automatically on GPUs with native support (RTX 30-series and newer, L4, A100), FP16 fallback otherwise.</p>
</div><div class="fusion-title title fusion-title-15 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">3. One honest bug fix: per-model output files</span></h2></div><div class="fusion-text fusion-text-25 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">While validating the new backend, we caught a fossil from UVLM&#8217;s two-backend era: the notebooks used a hardcoded rule that sent every non-Qwen2.5 model&#8217;s results to <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">Score_Analysis_LLaVA.csv</code>. With four families, that meant different models could silently append into the same CSV — and resume mode could not tell them apart. As of v3.2.0, <strong>output filenames are derived from the loaded checkpoint</strong> (e.g. <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">Score_Analysis_InternVL3_5-8B-HF.csv</code>), so each model writes its own file and resume mode and schema upgrades are per-model by construction. If you benchmark several models on the same image folder, this is the release that keeps your results honest.</p>
</div><div class="fusion-title title fusion-title-16 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">4. Getting started</span></h2></div><div class="fusion-text fusion-text-26 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:5px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Nothing changes in the workflow — install (or upgrade) and the new family is there:</p>
</div><div class="fusion-text fusion-text-27 fusion-text-no-margin" style="--awb-margin-top:5px;--awb-margin-bottom:5px;"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="Bash1" data-enlighter-title="Bash">pip install --upgrade --force-reinstall --no-deps
git+https://github.com/perezjoan/UVLM.git</pre>
</div><div class="fusion-text fusion-text-28 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:5px;--awb-margin-bottom:5px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Or open the Colab notebook — it always installs the latest version automatically. The three-block workflow (load → configure tasks → run batch), consensus validation, chain-of-thought mode, and truncation detection all work with InternVL3.5 out of the box. No dependency changes since v3.1.0. Tested locally on Windows 11 with an RTX 5060 laptop GPU, where the 1B model loads in about 15 seconds once cached.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">One field note from validation, and a nice illustration of why UVLM separates format reliability from accuracy: under temperature sampling, InternVL3.5 1B answered a counting task with &#8220;There are two vehicles in the picture&#8221; — correct, but unparseable as an integer, so it was recorded as NA by design. Under greedy decoding with a strict format instruction, the same model returned a clean integer. Small models follow instructions best when you ask firmly and decode greedily.</p>
</div><div class="fusion-title title fusion-title-17 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">5. What&#8217;s next</span></h2></div><div class="fusion-text fusion-text-29 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The third family addition promised in v3.1.0 — the <strong>Gemma</strong> multimodal line — is coming next, and it will be a bigger step than a minor version: Gemma 4 requires the Transformers v5 line, which means UVLM&#8217;s next release will be a <strong>major version</strong> with a documented migration. Same discipline as always: one backend at a time, validated before released.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Full change log in VERSIONS.txt · Source and releases on <a class="keychainify-checked" href="https://github.com/perezjoan/UVLM">GitHub</a> · If you use UVLM in research, please cite our <a class="keychainify-checked" href="https://www.mdpi.com/2674-113X/5/3/30">Software paper</a>.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-7 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-30"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--3" data-awb-toc-id="3" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-8 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/">UVLM v3.2.0 — InternVL3.5 Joins the Registry, With Zero Notebook Changes</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>The AI Reading Series · Lecture 2: How Connectionism Conquered Language</title>
		<link>https://urbangeoanalytics.com/from-neurons-to-llms-transformer-history/</link>
					<comments>https://urbangeoanalytics.com/from-neurons-to-llms-transformer-history/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Tue, 11 Aug 2026 11:28:52 +0000</pubDate>
				<category><![CDATA[Getting Started]]></category>
		<category><![CDATA[theory]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[lecture]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2928</guid>

					<description><![CDATA[<p>Part 2 of the AI reading series. The previous post ended with AlexNet's 2012 earthquake in image recognition. This one tells the road to language: how researchers turned words into vectors, taught networks to read sequences, discovered attention — and why, in 2017, eight Google researchers decided attention was all you need.</p>
<p>The post <a href="https://urbangeoanalytics.com/from-neurons-to-llms-transformer-history/">The AI Reading Series · Lecture 2: How Connectionism Conquered Language</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-5 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-8 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-9 hover-type-none"><img decoding="async" width="1804" height="673" title="ai reading lecture 2" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/ai-reading-lecture-2.png" alt class="img-responsive wp-image-2975" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/ai-reading-lecture-2-200x75.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/ai-reading-lecture-2-400x149.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/ai-reading-lecture-2-600x224.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/ai-reading-lecture-2-800x298.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/ai-reading-lecture-2-1200x448.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/ai-reading-lecture-2.png 1804w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-9 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-31"><p><em>Part 2 of the AI reading series. Read this after <a class="keychainify-checked" href="https://urbangeoanalytics.com/understanding-modern-ai-lecture-1-cardon-neurons-spike-back/">Part 1: Where Machine Learning Came From.</a></em></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="9:1-9:412;283-694">The paper covered in the previous post of this series ends at the moment neural networks triumph: 2012, image recognition, ImageNet, the AlexNet &#8220;earthquake.&#8221; It gave us a precious reading grid — the <strong>world</strong>, the <strong>calculator</strong>, the <strong>horizon</strong> — and one central idea: connectionist machines won by <em>emptying the calculator</em> of all explicit rules, letting enormous masses of data shape the program themselves.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="11:1-11:707;696-1402">But Cardon and his colleagues talk mostly about <strong>images</strong>. The paper you are about to read next, &#8220;<a class="keychainify-checked" href="https://arxiv.org/abs/1706.03762">Attention Is All You Need</a>&#8220;, is about <strong>language</strong> — machine translation, to be precise. And there is a world of difference between recognizing a rhinoceros in a photo and translating a sentence from French to English. This post tells the story of that road: how, between the mid-1980s and 2017, researchers adapted neural networks to the most &#8220;symbolic&#8221; problem there is — words, sentences, meaning — until they produced the architecture that now dominates all of artificial intelligence: the <strong>Transformer</strong>, the building block of <strong>Large Language Models</strong> (LLMs) such as ChatGPT, Claude, Gemini, and Qwen.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="13:1-13:164;1404-1567">We will start simple and ramp up the technicality progressively. At the end you will find a timeline and a glossary: keep them at hand while reading Vaswani et al.</p>
</div><div class="fusion-title title fusion-title-18 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">1. The problem images never posed: sequence</span></h2></div><div class="fusion-text fusion-text-32 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="19:1-19:436;1622-2057">Let&#8217;s pick up where the previous paper left off. In 2012, a convolutional neural network (CNN) crushes the ImageNet competition. Why does <strong>convolution</strong> work so well on images? Because an image is a <em>spatial</em> object: a rhinoceros is still a rhinoceros whether it stands on the left or the right of the photo. The CNN exploits this property by sliding small filters across the image, like a magnifying glass sweeping over a photograph.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="21:1-21:115;2059-2173">Language poses a different problem. A sentence is a <strong>sequence</strong>: a <em>temporal</em>, ordered object of variable length.</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="23:1-25:85;2175-2611">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="23:1-23:112;2175-2286">&#8220;The dog bites the man&#8221; and &#8220;The man bites the dog&#8221; contain the same words, but <strong>order changes everything</strong>.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="24:1-24:240;2287-2526">The meaning of a word can depend on words that are far away: in &#8220;The key that I left on the kitchen table last night <strong>has disappeared</strong>,&#8221; the verb agrees with &#8220;the key,&#8221; a dozen words earlier. This is called a <strong>long-range dependency</strong>.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="25:1-25:85;2527-2611">A sentence can be 3 words or 300: the network must accept inputs of variable size.</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="27:1-27:209;2613-2821">Remember the formula from the previous post: for connectionists, the goal is to &#8220;put the world into a vector.&#8221; Two questions immediately arise, and the whole story that follows is a series of answers to them:</p>
<ol class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-decimal flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="29:1-30:92;2823-2988">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="29:1-29:74;2823-2896"><strong>How do you turn a word into a vector?</strong> (the representation problem)</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="30:1-30:92;2897-2988"><strong>How do you process a variable-length sequence of vectors?</strong> (the architecture problem)</li>
</ol>
</div><div class="fusion-title title fusion-title-19 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">2. First answer: turning words into vectors</span></h2></div><div class="fusion-text fusion-text-33 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="36:1-36:52;3043-3094"><strong>The word as a number (and why it is not enough)</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="38:1-38:496;3096-3591">The naive approach gives each dictionary word a number: &#8220;cat&#8221; = 4,812, &#8220;dog&#8221; = 1,953. Technically, one uses a so-called <em>one-hot</em> vector: a huge vector of zeros with a single 1 at the word&#8217;s position. The problem? In this representation, &#8220;cat&#8221; is no closer to &#8220;dog&#8221; than to &#8220;umbrella.&#8221; All resemblance between words is lost. It is a <strong>symbolic</strong> representation in the sense of the previous post: each word is a discrete symbol, identical to or different from another, with no internal structure.</p>
<p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="40:1-40:45;3593-3637"><strong>Word2vec (2013): meaning as neighborhood</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="42:1-42:355;3639-3993">The solution was already hinted at in Cardon&#8217;s paper: <strong>word2vec</strong> (Mikolov et al., 2013). The idea rests on an old linguists&#8217; intuition, summed up by John R. Firth in 1957: <em>&#8220;You shall know a word by the company it keeps.&#8221;</em> &#8220;Cat&#8221; and &#8220;dog&#8221; appear in similar contexts (&#8220;feed the ___&#8221;, &#8220;the ___ is sleeping&#8221;); they should therefore receive nearby vectors.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="44:1-44:322;3995-4316">Word2vec trains a small neural network on a simple task: guess a word from its neighbors (or the reverse). The network is not the point; what you keep are the vectors learned along the way, called <strong>embeddings</strong>. Each word becomes a point in a space of a few hundred dimensions, and this space has astonishing properties:</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="46:1-47:119;4318-4493">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="46:1-46:57;4318-4374">words that are close in meaning are close in distance;</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="47:1-47:119;4375-4493">some directions of the space <em>mean</em> something: the famous computation <strong>king − man + woman ≈ queen</strong> actually works.</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="49:1-49:362;4495-4856">Now reread that sentence from Cardon&#8217;s paper: semantic proximity is not deduced from a symbolic categorization, but induced from statistical neighborhoods. This is exactly the symbolic-to-connectionist reversal, applied to the meaning of words. No linguist wrote a rule; meaning emerged from data. The first of our two problems — representing words — is solved.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="51:1-51:368;4858-5225">But word2vec has a serious limitation: each word gets <strong>one single, fixed vector</strong>. Yet &#8220;bank&#8221; does not mean the same thing in &#8220;I called my bank&#8221; and &#8220;I sat on the river bank.&#8221; What we would need are <strong>contextual</strong> representations that change with the sentence. Keep this limitation in mind: the Transformer is, among other things, the machine that will blow it away.</p>
</div><div class="fusion-title title fusion-title-20 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">3. Second answer: recurrent networks, a memory that reads word by word</span></h2></div><div class="fusion-text fusion-text-34 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="57:1-57:42;5307-5348"><strong>The RNN: a loop over time (1986–1990)</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="59:1-59:295;5350-5644">There remained the second problem: processing a variable-length sequence. The historical answer is the <strong>recurrent neural network</strong> (RNN), popularized by Jeffrey Elman in 1990 and already present in the work of the PDP group of Rumelhart and Hinton (1986) — the same group as in Cardon&#8217;s paper.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="61:1-61:258;5646-5903">The idea is elegant: the network reads the sequence <strong>one word at a time</strong>, left to right, and maintains a <strong>hidden state</strong> — a vector acting as working memory. At each word, the network combines what it reads with what it remembers, and updates its memory:</p>
<blockquote class="ml-2 border-l-4 border-&#091;hsl(var(--border-300)/0.1)&#093; pl-4 text-text-300" data-sourcepos="63:1-63:40;5905-5944">
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="63:3-63:40;5907-5944">memory(t) = f( memory(t−1), word(t) )</p>
</blockquote>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="65:1-65:229;5946-6174">Picture someone reading a sentence under their breath while keeping a mental summary they revise at every word. That is an RNN. The architecture naturally accepts sentences of any length: you simply loop for more or fewer steps.</p>
<p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="67:1-67:35;6176-6210"><strong>The vanishing gradient problem</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="69:1-69:540;6212-6751">In practice, simple RNNs have a crippling flaw. Remember <strong>backpropagation</strong>: to learn, the error is propagated backwards through the network. In an RNN, &#8220;backwards&#8221; means <strong>back through time</strong>, across as many steps as there are words. At each step, the error signal is multiplied by coefficients; over a long sentence it gets multiplied dozens of times by numbers that are often smaller than 1&#8230; and it melts like snow in the sun. This is the famous <strong>vanishing gradient</strong> problem, identified notably by Sepp Hochreiter as early as 1991.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="71:1-71:188;6753-6940">The concrete consequence: the network cannot learn long-range dependencies. By the time it reaches the verb &#8220;has disappeared,&#8221; it has &#8220;forgotten&#8221; the key at the beginning of the sentence.</p>
<p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="73:1-73:41;6942-6982"><strong>The LSTM (1997): a memory with gates</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="75:1-75:510;6984-7493">The most celebrated solution arrives in 1997: the <strong>LSTM</strong> (<em>Long Short-Term Memory</em>), by Sepp Hochreiter and Jürgen Schmidhuber. The idea: equip the memory cell with learned <strong>gates</strong> — small mechanisms that decide, at each step, what to <strong>write</strong> into memory, what to <strong>forget</strong>, and what to <strong>read</strong>. A kind of whiteboard with a doorkeeper choosing what gets noted and what gets erased. Thanks to an internal &#8220;conveyor belt&#8221; where information circulates almost untouched, the gradient survives much longer.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="77:1-77:331;7495-7825">LSTMs (and their simplified cousin, the <strong>GRU</strong>, 2014) would dominate natural language processing for twenty years. They are a perfect example of what Cardon&#8217;s paper calls the work on <strong>hyper-parameters</strong> and architecture: you do not dictate grammar rules to the network — you <em>sculpt</em> its structure so it can learn what it needs.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="79:1-79:321;7827-8147">Around 2014–2015, boosted by GPUs and large corpora, LSTMs power the speech recognition in your phone and the first neural versions of Google Translate. Connectionism, victorious over images in 2012, is winning over language. But it is precisely by pushing LSTMs to their limits that the missing link will be discovered.</p>
</div><div class="fusion-title title fusion-title-21 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">4. Translating: the seq2seq model and its bottleneck</span></h2></div><div class="fusion-text fusion-text-35 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="85:1-85:27;8211-8237"><strong>Encoder–decoder (2014)</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="87:1-87:313;8239-8551">Machine translation is THE queen of tasks, because it demands everything: understanding one sequence and producing another, of a different length. In 2014, two teams (Sutskever, Vinyals and Le at Google; Cho and Bengio in Montreal) propose the <strong>seq2seq</strong> (<em>sequence-to-sequence</em>) architecture, made of two RNNs:</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="89:1-90:181;8553-8918">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="89:1-89:185;8553-8737">an <strong>encoder</strong> reads the source sentence (&#8220;Le chat dort&#8221;) and compresses it into <strong>a single vector</strong> — a numerical summary of the whole sentence, sometimes called a &#8220;thought vector&#8221;;</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="90:1-90:181;8738-8918">a <strong>decoder</strong> starts from this vector and generates the target sentence word by word (&#8220;The,&#8221; then &#8220;cat,&#8221; then &#8220;sleeps&#8221;), each produced word being fed back in to produce the next.</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="92:1-92:232;8920-9151">Note this word-by-word generation mechanism, each word conditioned on the previous ones: it is called <strong>auto-regressive</strong>, and it is <em>exactly</em> how ChatGPT or Claude write their answers today. This point from 2014 has never changed.</p>
<p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="94:1-94:19;9153-9171"><strong>The bottleneck</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="96:1-96:364;9173-9536">But seq2seq has an obvious Achilles heel: the <strong>entire</strong> source sentence, whether 5 or 60 words long, must fit into a single fixed-size vector. It is like asking a translator to read a whole paragraph, close the book, then translate from memory without ever reopening it. On long sentences, quality collapses. Researchers call this the information <strong>bottleneck</strong>.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="98:1-98:68;9538-9605">The solution will give its name to the paper you are about to read.</p>
</div><div class="fusion-title title fusion-title-22 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">5. Attention: letting the network look wherever it wants</span></h2></div><div class="fusion-text fusion-text-36 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="104:1-104:244;9685-9928">In 2014, Dzmitry Bahdanau, Kyunghyun Cho and Yoshua Bengio (him again — find him in the connectionist core of Cardon&#8217;s figure 2) publish an idea that changes everything: what if, instead of closing the book, the decoder could <strong>keep it open</strong>?</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="106:1-106:550;9930-10479">Concretely: the encoder no longer produces a single vector, but keeps one vector <strong>per word</strong> of the source sentence. Then, for every word it generates, the decoder computes a <strong>relevance score</strong> between what it is currently doing and each of the source words. These scores, turned into percentages (via a function called <em>softmax</em>), are used to build a weighted average of the source vectors: the decoder &#8220;focuses&#8221; on the words that are useful at that instant. To produce &#8220;sleeps,&#8221; it looks mostly at &#8220;dort.&#8221; This mechanism is called <strong>attention</strong>.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="108:1-108:88;10481-10568">Three things to remember, because they directly prepare your reading of Vaswani et al.:</p>
<ol class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-decimal flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="110:1-112:223;10570-11193">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="110:1-110:266;10570-10835">Attention is <strong>learned</strong>, not programmed. Nobody wrote a French–English alignment rule: the network discovers by itself that &#8220;dort&#8221; explains &#8220;sleeps.&#8221; Once again, the inductive move described by Cardon — empty the calculator, let the world provide the structure.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="111:1-111:135;10836-10970">Attention is <strong>soft</strong>: it is not a binary choice but a weighting, therefore it is differentiable, therefore backprop applies to it.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="112:1-112:223;10971-11193">Attention creates <strong>shortcuts</strong>: every generated word is directly connected to every source word, without going through the fragile chain of recurrent memory. Long-range dependencies become connections&#8230; of distance 1.</li>
</ol>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="114:1-114:548;11195-11742">Between 2015 and 2017, the &#8220;LSTM + attention&#8221; recipe becomes the world state of the art in translation (Google deploys it at the end of 2016). But an irritation grows among engineers, and it is a <em>hardware</em> one — remember the role of GPUs in Cardon&#8217;s story. An RNN reads a sentence <strong>word after word</strong>: computing word 50 must wait for word 49. Yet GPUs are massively <strong>parallel</strong> machines, built to perform millions of operations <em>at the same time</em>. Recurrence wastes the hardware and forbids training on truly gigantic corpora in reasonable time.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="116:1-116:116;11744-11859">Hence a question asked ever more insistently in the labs: now that we have attention, what is recurrence still for?</p>
</div><div class="fusion-title title fusion-title-23 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">6. &#8220;Attention Is All You Need&#8221;</span></h2></div><div class="fusion-text fusion-text-37 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="122:1-122:384;11907-12290">The answer from eight Google researchers (Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin) fits in their slightly provocative title: <em>attention is all you need</em>. Their architecture, the <strong>Transformer</strong>, purely and simply removes recurrence (and convolution). Here, as a preview, are the ideas you will meet in the paper — consider this section your reading map.</p>
<p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="124:1-124:19;12292-12310"><strong>Self-attention</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="126:1-126:452;12312-12763">Until now, attention connected the decoder to the encoder (target to source). The Transformer&#8217;s stroke of genius is to apply it <strong>inside a single sentence</strong>: every word looks at every other word of its own sentence — including itself — to build its representation. In &#8220;The animal didn&#8217;t cross the street because <strong>it</strong> was tired,&#8221; the word &#8220;it&#8221; learns to attend to &#8220;the animal.&#8221; In one operation, each word enriches its meaning with its whole context.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="128:1-128:208;12765-12972">Notice what this solves: a word&#8217;s vector is no longer fixed as in word2vec — it is <strong>contextual</strong>, recomputed for every sentence. &#8220;Bank&#8221; will not have the same representation at the counter and by the river.</p>
<p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="130:1-130:44;12974-13017"><strong>Query, Key, Value: the library metaphor</strong></p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="132:1-132:560;13019-13578">The paper formalizes attention with three vectors per word, obtained through three learned transformations: a <strong>Query</strong>, a <strong>Key</strong>, and a <strong>Value</strong>. The classic image: in a library, you arrive with a question (your <em>query</em>), you compare it to the labels on the spines of the books (the <em>keys</em>), and you leave with the content of the relevant books (the <em>values</em>), in proportion to their relevance. Mathematically: dot products between Q and K → scores → softmax → weighted average of the V. That is equation (1) of the paper, the only truly indispensable one:</p>
<blockquote class="ml-2 border-l-4 border-&#091;hsl(var(--border-300)/0.1)&#093; pl-4 text-text-300" data-sourcepos="134:1-134:45;13580-13624">
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="134:3-134:45;13582-13624">Attention(Q, K, V) = softmax(QKᵀ / √d) · V</p>
</blockquote>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="136:1-136:166;13626-13791">The √d in the denominator is a mere numerical safeguard (keeping the scores from growing too large). Everything else in the paper is engineering around this formula.</p>
<p class="mt-2 -mb-1 text-base font-bold" dir="ltr" data-sourcepos="138:1-138:52;13793-13844"><strong>Multi-head, position, and the full architecture</strong></p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="140:1-143:298;13846-15137">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="140:1-140:256;13846-14101"><strong>Multi-head attention</strong>: rather than one attention, you run 8 (or 16, or 96&#8230;) in parallel, each with its own Q, K, V. Each &#8220;head&#8221; specializes — one tracks syntax, another coreference, and so on — without anyone asking it to. Emergent behavior, again.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="141:1-141:378;14102-14479"><strong>Positional encoding</strong>: removing recurrence has a cost — the network no longer knows in what order the words are! (Attention is a set computation, insensitive to order.) The fix: add to each embedding a small vector encoding its position, built from sines and cosines of varying frequencies. Do not drown in the formulas: just remember that we <em>re-inject</em> the order we lost.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="142:1-142:360;14480-14839"><strong>Stacking</strong>: one Transformer &#8220;block&#8221; = self-attention + a small classical network (<em>feed-forward</em>), with two training stabilizers (residual connections and normalization). These blocks are stacked: 6 in the 2017 paper, 96 and more in today&#8217;s LLMs. The original paper keeps the encoder–decoder structure inherited from seq2seq, since it targets translation.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="143:1-143:298;14840-15137"><strong>Parallelism</strong>: all words are processed <strong>at the same time</strong>. No more sequential waiting: the Transformer fits GPUs perfectly. This is the decisive argument — less training time, more data ingested. Reread Cardon: the connectionist victory has always been as much about hardware as about ideas.</li>
</ul>
</div><div class="fusion-title title fusion-title-24 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">7. From the Transformer to the LLM</span></h2></div><div class="fusion-text fusion-text-38 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="149:1-149:130;15215-15344">The 2017 paper is about translation. How do we get to ChatGPT? Through an idea of disarming simplicity: <strong>next-word prediction</strong>.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="151:1-151:426;15346-15771">Take a text, hide the next word, ask the model to guess it, correct it with backprop, repeat — trillions of times, over the whole web. No human labels are needed: the text is its own correction. This is called <strong>self-supervised</strong> learning. Think back to Cardon&#8217;s &#8220;horizon&#8221;: here, the horizon of the computation (the next word) is supplied by the world itself, for free, at unlimited scale. It is the ultimate cybernetic loop.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="153:1-153:20;15773-15792">The key milestones:</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="155:1-159:481;15794-17724">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="155:1-155:350;15794-16143"><strong>2018 — GPT-1</strong> (OpenAI): keep only the <strong>decoder</strong> of the Transformer, trained to predict the next word, then fine-tuned on specific tasks. <strong>BERT</strong> (Google) makes the opposite choice — keep only the <strong>encoder</strong>, trained to guess masked words — and crushes the comprehension benchmarks. The two descendants of the 2017 paper divide up the world.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="156:1-156:275;16144-16418"><strong>2019 — GPT-2</strong>: same recipe, ×10 on size (1.5 billion parameters). Surprise: the model can summarize, translate, answer questions <em>without having been trained to</em> — simply because predicting the next word across the whole Internet forces it to learn a bit of everything.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="157:1-157:396;16419-16814"><strong>2020 — GPT-3 and the scaling laws</strong> (Kaplan et al.): performance grows in a regular, <em>predictable</em> way with three ingredients — parameters, data, compute. The message: bigger is better, and nobody sees the ceiling yet. GPT-3 (175 billion parameters) reveals <strong>in-context learning</strong>: give it two examples inside the question (the <em>prompt</em>), and it picks up the pattern without any retraining.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="158:1-158:429;16815-17243"><strong>2022 — ChatGPT and RLHF</strong>: a raw LLM completes text; it does not &#8220;answer.&#8221; It is aligned in two stages: fine-tuning on dialogues written by humans, then <strong>reinforcement learning from human feedback</strong> (RLHF) — annotators rank answers, the model learns to aim for the best-ranked ones. Note the historical irony: the horizon of the computation, expelled from symbolic rules, comes back through the door of <em>human preferences</em>.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="159:1-159:481;17244-17724"><strong>2023–today</strong>: the era of open models (Meta&#8217;s Llama, Mistral, Alibaba&#8217;s <strong>Qwen</strong>), of multimodal models (text + image + sound), and of the Transformer&#8217;s extension to&#8230; image and video generation. Modern diffusion models (Stable Diffusion 3, Flux, Qwen-Image, Sora, Wan) have also replaced their old internal networks with Transformers (the so-called <strong>DiT</strong>, <em>Diffusion Transformer</em>, architecture). The 2017 architecture has become the universal architecture of connectionism.</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="161:1-161:806;17726-18531">One last thing, to close the loop with Cardon. The final table of his paper describes deep learning as: <em>world</em> = vectors of massive data, <em>calculator</em> = deep network, <em>horizon</em> = error optimization on an objective. The LLM is its most extreme culmination: the world is <strong>all the text ever written</strong>, the calculator is a Transformer with hundreds of billions of coefficients, and the horizon fits in three words — <strong>predict the next word</strong>. That capacities for reasoning, translation and dialogue <em>emerge</em> from such a poor objective is perhaps the most beautiful posthumous victory of the connectionist camp — and the philosophical question remains wide open, exactly where Cardon left it: what does a machine &#8220;understand&#8221; when its entire thought is, in Hinton&#8217;s phrase, &#8220;a big vector of neural activity&#8221;?</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="167:1-167:45;18585-18629">The paper is 11 dense pages. Reading advice:</p>
<ol class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-decimal flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr" data-sourcepos="169:1-172:138;18631-19387">
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="169:1-169:260;18631-18890"><strong>Read in this order</strong>: the abstract → the introduction (§1) → figure 1 (the architecture — keep it in view at all times) → §3.2 (attention, the heart of the paper) → the conclusion. The rest (§5–6, training details and translation results) can be skimmed.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="170:1-170:167;18891-19057"><strong>Do not get stuck on the math.</strong> Only one equation matters (Attention(Q,K,V), equation 1), and you already know its intuition: query → labels → weighted contents.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="171:1-171:192;19058-19249"><strong>Spot the words you now own</strong>: <em>recurrent</em>, <em>sequential computation</em>, <em>long-range dependencies</em>, <em>parallelizable</em> — you now know why they are there and what the paper is fighting against.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="172:1-172:138;19250-19387"><strong>A question to hold onto as you close the paper</strong>: why does removing recurrence make the addition of positional encoding <em>mandatory</em>?</li>
</ol>
</div><div class="fusion-title title fusion-title-25 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:20px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Timeline</span></h2></div>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Year</th>
<th align="left">Event</th>
<th align="left">Why it matters</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1943</td>
<td align="left"> Formal neuron (McCulloch &amp; Pitts)</td>
<td align="left">The atom of connectionism</td>
</tr>
<tr>
<td align="left">1957</td>
<td align="left">Perceptron (Rosenblatt)</td>
<td align="left">First learning machine</td>
</tr>
<tr>
<td align="left">1969</td>
<td align="left"><em>Perceptrons</em> (Minsky &amp; Papert)</td>
<td align="left">The &#8220;excommunication,&#8221; connectionist winter</td>
</tr>
<tr>
<td align="left">1986</td>
<td align="left">Backprop popularized (Rumelhart, Hinton, Williams)</td>
<td align="left">Deep networks become trainable</td>
</tr>
<tr>
<td align="left">1990</td>
<td align="left">Elman&#8217;s RNN</td>
<td align="left">The network that reads sequences</td>
</tr>
<tr>
<td align="left">1991/1997</td>
<td align="left">Vanishing gradient identified / LSTM (Hochreiter &amp; Schmidhuber)</td>
<td align="left">A memory that goes the distance</td>
</tr>
<tr>
<td align="left">2012</td>
<td align="left">AlexNet wins ImageNet</td>
<td align="left">The &#8220;earthquake&#8221; — where the previous lecture ends</td>
</tr>
<tr>
<td align="left">2013</td>
<td align="left">word2vec (Mikolov)</td>
<td align="left">Words become vectors of meaning</td>
</tr>
<tr>
<td align="left">2014</td>
<td align="left"> seq2seq (Sutskever; Cho)</td>
<td align="left">Encoder–decoder, auto-regressive generation</td>
</tr>
<tr>
<td align="left">2014–15</td>
<td align="left">Attention (Bahdanau, Cho, Bengio)</td>
<td align="left">The decoder keeps the book open</td>
</tr>
<tr>
<td align="left">2016</td>
<td align="left">Google Translate goes neural</td>
<td align="left">Connectionism conquers mainstream language</td>
</tr>
<tr>
<td align="left">2017</td>
<td align="left">&#8220;Attention Is All You Need&#8221; (Vaswani et al.)</td>
<td align="left">The Transformer: attention alone, parallel compute</td>
</tr>
<tr>
<td align="left">2018</td>
<td align="left">GPT-1 (decoder) and BERT (encoder)</td>
<td align="left">The two lineages of the Transformer</td>
</tr>
<tr>
<td align="left">2019</td>
<td align="left">GPT-2</td>
<td align="left">The &#8220;free&#8221; capabilities of next-word prediction</td>
</tr>
<tr>
<td align="left">2020</td>
<td align="left">GPT-3, scaling laws</td>
<td align="left"> Bigger = predictably better</td>
</tr>
<tr>
<td align="left">2022</td>
<td align="left">ChatGPT (RLHF)</td>
<td align="left">The LLM becomes a conversational assistant</td>
</tr>
<tr>
<td align="left">2023</td>
<td align="left">Llama, Mistral, Qwen, DiT</td>
<td align="left">Open models; the Transformer invades image/video diffusion</td>
</tr>
</tbody>
</table>
</div>
<div class="fusion-title title fusion-title-26 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Glossary</span></h2></div><div class="fusion-text fusion-text-39 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="202:1-202:205;20890-21094"><strong>Attention</strong> — A learned mechanism that computes, for a given element, relevance weights over a set of other elements, then takes their weighted average. Enables direct connections between distant words.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="204:1-204:161;21096-21256"><strong>Auto-regressive</strong> — Generation mode in which each new word is produced conditioned on all previous ones, one by one. This is how GPT, Claude and Qwen &#8220;write.&#8221;</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="206:1-206:123;21258-21380"><strong>Backpropagation</strong> — The algorithm (1986) that propagates the error from output to input to adjust the network&#8217;s weights.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="208:1-208:129;21382-21510"><strong>BERT</strong> (2018) — Encoder-only model, trained to guess masked words; excellent at <em>understanding</em> text (classification, search).</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="210:1-210:115;21512-21626"><strong>Context window</strong> — The maximum number of tokens the model can consider at once. A major practical limit of LLMs.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="212:1-212:124;21628-21751"><strong>Decoder</strong> — The half of the Transformer that <em>generates</em> the output sequence. GPT and most current LLMs are decoder-only.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="214:1-214:147;21753-21899"><strong>Embedding</strong> — Representation of an object (word, image, molecule&#8230;) as a dense vector in a space where geometric proximity reflects similarity.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="216:1-216:90;21901-21990"><strong>Encoder</strong> — The half of the Transformer that <em>reads</em> and represents the input sequence.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="218:1-218:105;21992-22096"><strong>GRU / LSTM</strong> — Gated recurrent cells (1997 for the LSTM) that protect information over long stretches.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="220:1-220:150;22098-22247"><strong>Hyper-parameters</strong> — Architecture and training choices set by humans (number of layers, heads, learning rate&#8230;), as opposed to learned parameters.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="222:1-222:154;22249-22402"><strong>In-context learning</strong> — A model&#8217;s ability (from GPT-3 onwards) to pick up a task from a few examples placed directly in the prompt, without retraining.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="224:1-224:151;22404-22554"><strong>LLM (Large Language Model)</strong> — A Transformer (usually decoder-only) with billions of parameters, trained by next-word prediction on immense corpora.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="226:1-226:157;22556-22712"><strong>Long-range dependency</strong> — A grammatical or semantic link between words that are far apart in a sentence. Achilles heel of RNNs, strong point of attention.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="228:1-228:138;22714-22851"><strong>Multi-head attention</strong> — Running several independent attentions (&#8220;heads&#8221;) in parallel, each free to specialize in one type of relation.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="230:1-230:160;22853-23012"><strong>Parameters (weights)</strong> — The coefficients adjusted during learning. GPT-3: 175 billion. These are what you download when you fetch a model from Hugging Face.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="232:1-232:114;23014-23127"><strong>Positional encoding</strong> — Vectors added to the embeddings to re-inject word order, which attention alone ignores.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="234:1-234:123;23129-23251"><strong>Prompt</strong> — The input text given to the LLM; since GPT-3, you &#8220;program&#8221; the model in natural language through the prompt.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="236:1-236:177;23253-23429"><strong>Query / Key / Value (Q, K, V)</strong> — The three learned projections of each word used in the attention computation: the question asked, the label compared, the content retrieved.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="238:1-238:142;23431-23572"><strong>RLHF</strong> — Reinforcement Learning from Human Feedback: aligning an LLM with human preferences via reinforcement (the basis of ChatGPT, 2022).</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="240:1-240:164;23574-23737"><strong>RNN (recurrent neural network)</strong> — A network that processes a sequence step by step while maintaining a hidden state (memory). Dominant in NLP from 1990 to 2017.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="242:1-242:131;23739-23869"><strong>Scaling laws</strong> — Empirical relations (2020) showing that LLM performance improves predictably with model size, data and compute.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="244:1-244:128;23871-23998"><strong>Self-supervised learning</strong> — Learning without human labels: the data provides its own target (e.g., the next word of a text).</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="246:1-246:169;24000-24168"><strong>Seq2seq</strong> — The encoder–decoder architecture (2014) turning one sequence into another; the framework of neural translation and the direct ancestor of the Transformer.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="248:1-248:237;24170-24406"><strong>Softmax</strong> — The function that converts a list of scores into a probability distribution (percentages summing to 100%). Used inside attention and to pick the next word — it is on this function that the <em>temperature</em> parameter operates.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="250:1-250:110;24408-24517"><strong>Token</strong> — The elementary unit of text a model manipulates (often a word fragment, ~¾ of a word on average).</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="252:1-252:190;24519-24708"><strong>Transformer</strong> (2017) — Architecture based solely on attention (no recurrence, no convolution), massively parallelizable; the foundation of all LLMs and, by now, of diffusion models (DiT).</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="254:1-254:161;24710-24870"><strong>Word2vec</strong> (2013) — A method producing static word embeddings from their contexts of occurrence; the ancestor of the Transformer&#8217;s contextual representations.</p>
</div><div class="fusion-title title fusion-title-27 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:-10px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Watch After Reading</span></h2></div><div class="fusion-text fusion-text-40" style="--awb-content-alignment:justify;"><ul>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="260:1-260:199;24901-25099"><strong>3Blue1Brown — the Transformer chapters of the Deep Learning series</strong>: &#8220;<a class="keychainify-checked" href="https://www.youtube.com/watch?v=wjZofJX0v4M">But what is a GPT?</a>&#8221; then &#8220;<a class="keychainify-checked" href="https://www.youtube.com/watch?v=eMlx5fFNoYc">Attention in transformers, step-by-step.</a>&#8221; The finest visualization of equation (1) in existence.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2" data-sourcepos="261:1-261:198;25100-25297"><strong>&#8220;<a class="keychainify-checked" href="https://www.youtube.com/watch?v=kCc8FmEb1nY">Let&#8217;s build GPT: from scratch, in code, spelled out</a>&#8221; — Andrej Karpathy</strong> (~2 h): building a GPT by following the paper, from tokenization to multi-head attention. Watch it with the code open.</li>
</ul>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-10 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-41"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--4" data-awb-toc-id="4" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-10 hover-type-zoomout"><img decoding="async" width="1536" height="1024" title="blog lvl1" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1.png" alt class="img-responsive wp-image-1685" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-6 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"></div></div></p>
<p>The post <a href="https://urbangeoanalytics.com/from-neurons-to-llms-transformer-history/">The AI Reading Series · Lecture 2: How Connectionism Conquered Language</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/from-neurons-to-llms-transformer-history/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>UVLM v3.1.0 — Qwen3-VL Joins the Registry, With Family-Based Model Selection</title>
		<link>https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/</link>
					<comments>https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Mon, 10 Aug 2026 07:46:22 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Package]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[Qwen]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2898</guid>

					<description><![CDATA[<p>UVLM v3.1.0 adds a third model family, Qwen3-VL (2B–32B Instruct), bringing the registry to 15 checkpoints. The notebooks gain a two-level family/model selector, the loader picks BF16 automatically on capable GPUs, and the smallest new model runs in about 2 GB of VRAM. Same three-block workflow, same prompts, one more family to compare.</p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/">UVLM v3.1.0 — Qwen3-VL Joins the Registry, With Family-Based Model Selection</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-7 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-11 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-42"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-43" style="--awb-margin-top:-30px;"><ul>
<li><strong data-start="64" data-end="88">New model family:</strong>Qwen3-VL (released from September 2025) joins LLaVA-NeXT and Qwen2.5-VL</li>
<li><strong>Two-level model selection</strong>: pick the family first, the model list refreshes automatically</li>
<li><strong>Lightest model yet</strong>: Qwen3-VL 2B runs in ~2 GB of VRAM with 4-bit quantization</li>
</ul>
</div><div class="fusion-title title fusion-title-28 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What is Qwen3-VL?</span></h2></div><div class="fusion-text fusion-text-44 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>UVLM was built around one idea: compare Vision-Language Models across architectures using <strong>identical prompts and evaluation protocols</strong>, without writing model-specific code. Until now that meant two families: LLaVA-NeXT and Qwen2.5-VL. Version 3.1.0 adds a third: <strong>Qwen3-VL</strong>, the successor to the Qwen2.5-VL family that anchored our published benchmark.</p>
<p>Qwen3-VL is the latest vision-language generation from Alibaba&#8217;s Qwen team, first released in <strong>September 2025</strong> with the 235B-A22B flagship, followed shortly after by the compact dense checkpoints that matter for most research budgets. UVLM v3.1.0 integrates the four dense Instruct sizes:</p>
</div>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Parameters</th>
<th align="left"> VRAM (4-bit)</th>
<th align="left">Typical hardware</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Qwen3-VL 2B Instruct</td>
<td align="left">2B</td>
<td align="left"> ~2 GB</td>
<td align="left">Any modern laptop GPU, free Colab T4</td>
</tr>
<tr>
<td align="left">Qwen3-VL 4B Instruct</td>
<td align="left">4B</td>
<td align="left"> ~3 GB</td>
<td align="left"> T4, RTX 3060</td>
</tr>
<tr>
<td align="left">Qwen3-VL 8B Instruct</td>
<td align="left">8B</td>
<td align="left"> ~6 GB</td>
<td align="left"> T4, RTX 4060/5060</td>
</tr>
<tr>
<td align="left">Qwen3-VL 32B Instruct</td>
<td align="left">32B</td>
<td align="left"> ~20 GB</td>
<td align="left"> A100, RTX 4090</td>
</tr>
</tbody>
</table>
</div>
<div class="fusion-text fusion-text-45 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">The registry now totals <strong>15 checkpoints across 3 families</strong>, from 2B to 110B parameters — and the 2B entry is the smallest model UVLM has ever supported, which makes it an interesting new baseline for large-scale, low-cost batch analysis.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="34:1-34:532;2755-3286">Technically, Qwen3-VL keeps the Qwen inference conventions (chat template → separate vision preprocessing → generation → token trimming), so it plugs into UVLM&#8217;s existing Qwen pipeline. What changes under the hood: the model loads through Transformers&#8217; generic <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">AutoModelForImageTextToText</code> class, requires <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">transformers ≥ 4.57</code> and <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">qwen-vl-utils ≥ 0.0.14</code>, and resizes images to multiples of 32 pixels rather than 28. All of this is handled inside the package — from the user&#8217;s side, it is simply one more family in the dropdown.</p>
</div><div class="fusion-title title fusion-title-29 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Pick the family, then the model</span></h2></div><div class="fusion-text fusion-text-46 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">With three families and fifteen checkpoints, a single flat dropdown was getting crowded. Both notebooks (Colab and local) now use a <strong>two-level selector</strong>: choose the family first — LLaVA-NeXT, Qwen2.5-VL, or Qwen3-VL — and the model list refreshes automatically.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-11" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-11 hover-type-none"><img decoding="async" width="1140" height="454" title="uvlm family" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family.png" alt class="img-responsive wp-image-2900" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-200x80.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-400x159.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-600x239.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-800x319.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family.png 1140w" sizes="(max-width: 640px) 100vw, 1140px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">UVLM v3.1.0: Two-level model selection</div></div></div></div><div class="fusion-text fusion-text-47 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">The selector is built from a new <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">FAMILY_GROUPS</code> mapping in the registry, which means future families will appear in the widgets automatically, with no notebook edits at all.</p>
</div><div class="fusion-title title fusion-title-30 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Smarter precision handling</span></h2></div><div class="fusion-text fusion-text-48 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">Qwen3-VL checkpoints are trained in BF16. On GPUs with native BF16 support (RTX 30-series and newer, L4, A100), the loader now selects <strong>BF16 automatically</strong>, falling back to FP16 on older cards and FP32 on CPU. If you followed our earlier benchmark work, you may remember the FP16 numerical-overflow crashes we documented with BF16-trained checkpoints on T4 hardware — this release is the first step toward closing that class of problem at the loader level.</p>
</div><div class="fusion-title title fusion-title-31 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Quality-of-life fixes</span></h2></div><div class="fusion-text fusion-text-49 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">Local Jupyter users get a long-overdue improvement: the model-loading progress (download bars, device map, timings) is now displayed in a <strong>log area under the Load button</strong>. Previously, output emitted inside the widget callback was silently swallowed in local Jupyter — Colab was never affected. The release also silences the <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">torch_dtype</code> deprecation warnings from recent Transformers versions and synchronizes the package version metadata.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="54:1-54:84;4795-4878">Nothing changes in the workflow — install (or upgrade) and the new family is there:</p>
</div><div class="fusion-text fusion-text-50 fusion-text-no-margin" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="bash1" data-enlighter-title="bash">pip install --upgrade --force-reinstall --no-deps git+https://github.com/perezjoan/UVLM.git</pre>
</div><div class="fusion-text fusion-text-51 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">Or open the <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://colab.research.google.com/github/perezjoan/UVLM/blob/main/notebooks/UVLM_colab.ipynb">Colab notebook</a> — it always installs the latest version automatically. The three-block workflow (load → configure tasks → run batch), consensus validation, chain-of-thought mode, and truncation detection all work with Qwen3-VL out of the box. Tested locally on Windows 11 with an RTX 5060 laptop GPU, where the 2B model loads in well under a minute once cached.</p>
</div><div class="fusion-title title fusion-title-32 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What’s next</span></h2></div><div class="fusion-text fusion-text-52 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="64:1-64:266;5471-5736">v3.1.0 is the first of a series of family additions. Next on the roadmap: <strong>InternVL3.5</strong> (via the Transformers-native <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">-HF</code> checkpoints) and the <strong>Gemma</strong> multimodal line. Each family will land as its own validated release — same discipline, one backend at a time.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="66:1-66:265;5738-6002">Full change log in <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://github.com/perezjoan/UVLM/blob/main/VERSIONS.txt">VERSIONS.txt</a> · Source and releases on <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://github.com/perezjoan/UVLM">GitHub</a> · If you use UVLM in research, please cite our <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://www.mdpi.com/2674-113X/5/3/30">Software paper</a>.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-12 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-53"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--5" data-awb-toc-id="5" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-12 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/">UVLM v3.1.0 — Qwen3-VL Joins the Registry, With Family-Based Model Selection</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>The AI Reading Series · Lecture 1: Where Machine Learning Came From</title>
		<link>https://urbangeoanalytics.com/understanding-modern-ai-lecture-1-cardon-neurons-spike-back/</link>
					<comments>https://urbangeoanalytics.com/understanding-modern-ai-lecture-1-cardon-neurons-spike-back/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Fri, 24 Jul 2026 02:53:32 +0000</pubDate>
				<category><![CDATA[Getting Started]]></category>
		<category><![CDATA[theory]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[lecture]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2540</guid>

					<description><![CDATA[<p>The first in a reading series taking you from the artificial neuron of 1943 to today's transformers, mixture-of-experts models and agents. Lecture 1 is a preparatory guide to Cardon, Cointet and Mazières' sociological history of AI, with reading strategy and glossary.</p>
<p>The post <a href="https://urbangeoanalytics.com/understanding-modern-ai-lecture-1-cardon-neurons-spike-back/">The AI Reading Series · Lecture 1: Where Machine Learning Came From</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-8 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-13 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-13 hover-type-none"><img decoding="async" width="1802" height="872" title="lecture 1 &#8211; long illustration" src="https://urbangeoanalytics.com/wp-content/uploads/2026/07/lecture-1-long-illustration.png" alt class="img-responsive wp-image-2546" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/07/lecture-1-long-illustration-200x97.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/07/lecture-1-long-illustration-400x194.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/07/lecture-1-long-illustration-600x290.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/07/lecture-1-long-illustration-800x387.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/07/lecture-1-long-illustration-1200x581.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/07/lecture-1-long-illustration.png 1802w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-14 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-54 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">This is the first post in a series of lectures intended to bring a reader with no formal background in machine learning up to the state of the field as it stands today. The trajectory runs from the artificial neuron of 1943 to the systems currently deployed in research and industry: deep convolutional networks, word and image embeddings, transformer architectures, large language models, vision-language models, mixture-of-experts routing, retrieval-augmented generation, and autonomous agents. Each lecture pairs a primary reading with a preparatory guide of this kind, plus a glossary and a set of questions to keep in mind while reading.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The series begins somewhere unexpected: <a class="keychainify-checked" href="https://hal.science/hal-02190026v1/file/NeuronsSpikeBack.pdf">a forty-page article written by sociologists</a>. This is deliberate. Before diving into LLM and agents, it is worth understanding where all of it came from, and above all understanding that none of these techniques was obvious or inevitable. Modern artificial intelligence is the outcome of a seventy-year scientific battle, with winners, losers, public humiliations, wilderness years, and spectacular comebacks. That is the story told by Dominique Cardon, Jean-Philippe Cointet and Antoine Mazières.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The article opens on a scene that reads like a western: October 2012, a scientific conference, a largely unknown student walks on stage and announces a result that pulverises ten years of work by an entire research community. The room is stunned. That scene, the earthquake of 2012, is the article&#8217;s destination rather than its starting point. Everything else explains how the field arrived there.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The thesis in one sentence: the history of AI is a war between two visions of the intelligent machine, a machine that is given rules and a machine that learns from examples, and after fifty years of symbolic domination it is the connectionists, long mocked and marginalised, who won.</p>
</div><div class="fusion-title title fusion-title-33 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">The Two Camps</span></h2></div><div class="fusion-text fusion-text-55 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Suppose you want to build a machine that recognises cats in photographs.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The symbolic approach says: sit down and write the rules. A cat has two pointed ears, whiskers, four legs; IF pointed ears AND whiskers THEN cat. This is intelligence as reasoning, the manipulation of symbols such as &#8220;ear&#8221; and &#8220;cat&#8221; by means of logic. It is intuitive, explainable and elegant, and it is how computers have always been programmed: the human writes the program, the machine executes it.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The connectionist approach says: rules are hopeless. A cat seen from behind, at night, half hidden behind a curtain, matches no rule at all. Instead, show the machine a hundred thousand photographs labelled &#8220;cat&#8221; or &#8220;not cat&#8221;, and let a network of small interconnected computing units, artificial neurons very loosely inspired by the brain, adjust the strength of its own connections until its answers are good. Nobody writes a rule; the rule emerges from the examples. This is intelligence as learning.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The strength of the article is that it shows this technical choice to be simultaneously a philosophical choice (is thinking reasoning, or perceiving?), an economic choice (who receives the funding?) and a social choice (which research communities dominate?).</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The article&#8217;s central image, figure 1, is worth studying closely:</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr">
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Classical machine, hypothetico-deductive:</strong> inputs plus program produce outputs. The human supplies the program.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Inductive machine:</strong> inputs plus outputs produce the program. The human supplies examples, photographs paired with correct answers, and the program is what comes out of the machine.</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">This reversal is the single most important idea in the entire series. If you retain one thing from Lecture 1, retain that one.</p>
</div><div class="fusion-title title fusion-title-34 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">The Authors’ Analytical Grid: World, Calculator, Horizon</span></h2></div><div class="fusion-text fusion-text-56 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The authors analyse each era of AI using three notions, announced early and reused all the way to the final synthesis table (table 1, page 28, an excellent summary of the whole article):</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr">
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>The world:</strong> what enters the machine. Data? Rules? Expert knowledge? An environment?</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>The calculator:</strong> what does the processing. A logic engine? A neural network? A black box?</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>The horizon:</strong> the goal of the computation. Solving a problem? Minimising an error? Imitating examples?</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Their key formula is slightly cryptic on first reading but becomes clear by the end. The symbolists wanted to put everything into the calculator, both the world and the goal, whereas the connectionists empty the calculator so that the world gives itself its own horizon. In other words, massive data supplies both the material and the correction: the labelled examples are what tell the network when it is wrong.</p>
</div><div class="fusion-title title fusion-title-35 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">The Four Eras and the Cast</span></h2></div><div class="fusion-text fusion-text-57 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The article follows four major periods. The characters introduced here reappear throughout the series.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Cybernetics and the first connectionism, 1943 to 1969.</strong> The age of the pioneers. Warren McCulloch and Walter Pitts invent the artificial neuron in 1943. Norbert Wiener founds cybernetics, the science of machines that self-correct through feedback. Frank Rosenblatt builds the Perceptron in 1957, the first machine that learns to recognise patterns; the press goes wild and conscious machines are promised.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Symbolic AI, 1956 to 1970, then expert systems in the 1980s.</strong> In 1956 John McCarthy and Marvin Minsky coin the term &#8220;artificial intelligence&#8221;, explicitly against cybernetics. With Herbert Simon and Allen Newell they capture the bulk of military funding and impose the symbolic vision. In 1969 Minsky publishes a book that proves neural networks have no future; this is the excommunication. Funding dries up, Rosenblatt dies in 1971, and connectionism enters a long winter. In the 1980s symbolic AI enjoys a second wind with expert systems, thousands of IF-THEN rules extracted from human doctors, geologists and engineers, before a second collapse.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>The return of the neurons, 1986 to 2010.</strong> A small group of holdouts, Geoffrey Hinton, Yann LeCun and Yoshua Bengio, later nicknamed the neural conspiracy, keeps the flame alive. In 1986 backpropagation finally makes it possible to train multi-layer networks. In 1989 LeCun gets a network to read postal codes, the first industrial application. But the years 1995 to 2007 are a colossal winter of rejected papers, mockery and isolation. The article contains excellent first-hand testimony from French researchers about that period.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>The triumph of deep learning, 2010 onward.</strong> Three ingredients converge: massive data, meaning the web and ImageNet with its fourteen million hand-labelled images; GPUs, the graphics processors built for video games and perfectly suited to the massively parallel computations of neural networks; and the algorithms patiently matured during the winter. Then 2012, and the earthquake. Since then, domain after domain, image, speech and text, deep networks have swept everything away.</p>
</div><div class="fusion-title title fusion-title-36 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">How to Read the Article Without Getting Lost</span></h2></div><div class="fusion-text fusion-text-58 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">It is an academic sociology article: dense, but well written and full of anecdotes. Some practical advice.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Take your time with the opening, the 2012 narrative on pages 2 and 3. The entire article is contained in that scene.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Lean on the figures. Figure 1, the two machines, and table 1, the four ages, are your two anchors. Figure 3, the timeline, shows the shifting dominations visually. Figures 4 and 5 give a first look at an artificial neuron and at backpropagation.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Do not get stuck on the details of the Web of Science queries in notes 5 and 6, or on the philosophical references to Fodor, Smolensky and the computational theory of mind. Grasp the general idea and move on. The section on convexity is the most technical; the glossary below gives the minimum needed to get through it.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Savour the interview quotations, set in indented blocks in spoken register. That is where the history comes alive.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Keep three questions in mind while reading:</p>
<ol class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-decimal flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr">
<li class="font-claude-response-body whitespace-normal break-words pl-2">For each era, what is in the world, the calculator and the horizon?</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2">Why did the connectionists win when they did, rather than in 1960? The hint is that it was not the ideas that changed.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2">The title speaks of striking back. Who was humiliated, when, and by whom?</li>
</ol>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Budget two to two and a half hours of attentive reading. It is the longest reading of the series and probably the one that will stay with you.</p>
</div><div class="fusion-title title fusion-title-37 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Survival Glossary</span></h2></div><div class="fusion-text fusion-text-59 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Terms are listed roughly in their order of appearance in the article.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Computer vision.</strong> Research field aiming to make machines see: recognising objects, faces and scenes in images.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>ImageNet.</strong> A database of fourteen million images across roughly twenty-one thousand categories, hand-labelled by thousands of micro-workers. It serves as an annual competition, and it is the benchmark on which the 2012 earthquake took place.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Benchmark.</strong> A standardised test set allowing objective comparison between the performance of different methods.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>GPU, graphics processing unit.</strong> A processor originally designed for video games, capable of performing millions of simple operations in parallel, which is exactly what a neural network requires.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Deep learning.</strong> Neural networks with many layers, hence deep. The term was coined by Hinton in 2006, partly to escape the poor reputation of the word connectionism.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Machine learning.</strong> The family of methods in which a machine learns from examples instead of being explicitly programmed. Deep learning is one branch of it.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Parameters, also weights or coefficients.</strong> The numbers adjusted during learning, encoding the strength of the connections between neurons. A hundred million parameters means a hundred million small knobs tuned automatically.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Symbolic AI, also GOFAI, Good Old-Fashioned AI.</strong> The rules-and-logic approach. Thinking equals manipulating symbols.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Connectionism.</strong> The neural network approach. Thinking equals parallel, distributed computation by simple units, with intelligence emerging from the connections.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Cybernetics.</strong> The science founded by Norbert Wiener in 1948, studying systems, whether machines or organisms, that regulate themselves through feedback. The direct ancestor of connectionism.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Feedback.</strong> Reinjecting the measured output error as a new input so that the system corrects itself. The thermostat is the canonical example. It is the founding principle of cybernetics and, in a sense, of all machine learning.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Black box.</strong> A system whose inputs and outputs are observable but whose internal workings are not understood. A recurring criticism of neural networks, and a property their defenders have claimed proudly since the cybernetic era.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Formal neuron.</strong> An ultra-simplified mathematical model of a neuron, due to McCulloch and Pitts in 1943. It sums its inputs weighted by weights and activates if the sum exceeds a threshold. Figure 4 shows that this amounts to three operations, no more.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Perceptron.</strong> The first learning machine built on a neural network, developed by Rosenblatt between 1957 and 1961, funded by the US Navy and designed for image recognition. It is the emblem of early connectionism and the target of Minsky&#8217;s attack.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Layers and hidden layers.</strong> Neurons are organised in tiers: an input layer holding the data, intermediate layers called hidden where the work happens, and an output layer holding the answer. Minsky&#8217;s 1969 book attacked a single-layer perceptron.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>XOR, exclusive OR.</strong> The elementary logical function &#8220;A or B, but not both&#8221;, which a single-layer perceptron cannot learn. This was Minsky and Papert&#8217;s decisive argument for burying connectionism. Multiple layers solve the problem.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>AI winter.</strong> A period of collapse in funding and credibility following excessive promises. There have been two general ones, in the early 1970s and the late 1980s, plus the specifically connectionist winters of 1969 to 1986 and 1995 to 2007.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Expert system.</strong> A 1980s program encoding human expert knowledge as thousands of IF-THEN rules, MYCIN for medical diagnosis being the standard example, driven by an inference engine that decides which rule to apply when. Expert systems mark both the peak and the collapse of symbolic AI.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Knowledge base.</strong> The stock of rules and facts held by an expert system, to be contrasted with a dataset: a knowledge base contains intelligible rules, a dataset contains raw examples.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Backpropagation, or backprop.</strong> The algorithm popularised in 1986 by Rumelhart, Hinton and Williams that makes learning possible. The error is measured at the output, then propagated backwards layer by layer to adjust each weight in the right direction. See figure 5. This is the central mechanism of all deep learning.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Loss function.</strong> The number measuring how wrong the network is. All of learning consists in minimising it.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Gradient descent.</strong> The minimisation method: compute the slope of the loss function and take a small step downhill, then repeat millions of times. The usual image is walking down a mountain in fog by following the slope underfoot. Stochastic means the slope is estimated on a small sample of the data at a time, which is faster.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Convolution and convolutional networks, CNNs.</strong> A technique invented by LeCun in 1989 for images. Rather than connecting every pixel to every neuron, small filters are slid across the image to detect local patterns such as edges and corners independently of their position. This is the architecture behind AlexNet in 2012.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Feature engineering.</strong> The now largely extinct art of hand-programming the relevant characteristics of the data, detecting edges, corners and contrasts, before passing them to an algorithm. Deep learning made it obsolete because the network discovers its own features. An entire scientific community lost its object of research this way, which accounts for the bitterness audible in some of the testimony in the article.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>End-to-end.</strong> Processing raw data, the pixels, all the way to the final answer, &#8220;cat&#8221;, within a single network, with no intermediate step programmed by a human.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>SVM, support vector machines, and kernel methods.</strong> A rival learning method dating from 1992, mathematically elegant and dominant between roughly 1995 and 2010. It was the great internal adversary within the learning camp, and the duel between SVMs and neural networks structures an entire section of the article.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Convexity.</strong> A mathematical criticism levelled at neural networks. A convex function is shaped like a bowl, with a single hollow, so gradient descent is guaranteed to reach the bottom, the global minimum. A neural network&#8217;s loss function is instead a landscape of mountains with countless valleys, local minima, with no guarantee of finding the best one. The SVM mathematicians treated this as a fatal flaw; LeCun&#8217;s answer amounted to saying that the theoretical guarantee matters less than the fact that it works better in practice, and experience proved him right.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Overfitting.</strong> When a network learns its training examples by heart instead of extracting general regularities, performing excellently on known data and poorly on new data. Dropout, the random switching-off of neurons during training, is one countermeasure mentioned in the article.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Hyper-parameters.</strong> All the architectural choices fixed by the human before learning begins: the number of layers, the number of neurons, the learning rate. These stand in contrast to parameters, which are learned automatically. The article shows that human labour does not disappear, it shifts from writing rules to setting these values.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Dataset.</strong> A collection of examples used to train and test a model, often as input-output pairs such as a photograph and its label.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Labelled data.</strong> Examples accompanied by the correct answer, supplied by humans. This is the fuel of supervised learning.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Crowdsourcing and Mechanical Turk.</strong> Micro-work platforms where thousands of people, paid by the task, label data by drawing a box around the dog or typing the spoken word. This is the human face, invisible and poorly paid, of so-called raw data, and ImageNet is its product.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Embedding, or vector.</strong> The transformation of an object, a word, an image, a social network, into a list of numbers so that a network can compute with it. The article cites word2vec and LeCun&#8217;s formula for putting the world into a vector, world2vec. This is the direct bridge to Lecture 2.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><strong>Induction and deduction.</strong> Deduction starts from general rules and applies them to particular cases, which is the symbolic approach. Induction starts from particular cases, the examples, and extracts a general rule, which is the connectionist approach. The article&#8217;s full subtitle, &#8220;the invention of inductive machines&#8221;, says the essential.</p>
</div><div class="fusion-title title fusion-title-38 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">After the Reading: Videos</span></h2></div><div class="fusion-text fusion-text-60 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Watch these after finishing the article to go further.</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3 print:block print:space-y-1" dir="ltr">
<li class="font-claude-response-body whitespace-normal break-words pl-2"><a class="keychainify-checked" href="https://www.youtube.com/watch?v=UZDiGooFs54"><em>The moment we stopped understanding AI [AlexNet]</em></a>, Welch Labs, about eighteen minutes, in English. The 2012 earthquake seen from inside the network, with excellent visualisations of AlexNet&#8217;s layers, and the perfect complement to the article&#8217;s opening scene.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><em>Heroes of Deep Learning: Andrew Ng interviews Geoffrey Hinton</em>, followed by the LeCun interview in the same series, in English. Both are cited in the article&#8217;s own footnotes, notes 3 and 23. Hearing Hinton and LeCun recount the connectionist winter in their own voices is worth the time.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><a class="keychainify-checked" href="https://www.youtube.com/watch?v=trWrEWfhTVg"><em>Le deep learning</em></a>, ScienceEtonnante (David Louapre), in French. Neural networks, backpropagation and convolution explained in twenty minutes.</li>
</ul>
</div><div class="fusion-title title fusion-title-39 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Reference</span></h2></div><div class="fusion-text fusion-text-61 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Cardon, D., Cointet, J.-P. and Mazières, A. (2018). <a class="keychainify-checked" href="https://hal.science/hal-02190026v1/file/NeuronsSpikeBack.pdf"><em>Neurons spike back. The invention of inductive machines and the artificial intelligence controversy.</em></a> Réseaux, 211(5), 173–220.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-15 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-62"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--6" data-awb-toc-id="6" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-14 hover-type-zoomout"><img decoding="async" width="1536" height="1024" title="blog lvl1" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1.png" alt class="img-responsive wp-image-1685" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl1.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-9 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"></div></div></p>
<p>The post <a href="https://urbangeoanalytics.com/understanding-modern-ai-lecture-1-cardon-neurons-spike-back/">The AI Reading Series · Lecture 1: Where Machine Learning Came From</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/understanding-modern-ai-lecture-1-cardon-neurons-spike-back/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Deploy Your Own Local LLM on Low VRAM in 30 Minutes — A Private Chat Assistant in Jupyter</title>
		<link>https://urbangeoanalytics.com/deploy-local-llm-low-vram-jupyter/</link>
					<comments>https://urbangeoanalytics.com/deploy-local-llm-low-vram-jupyter/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Tue, 02 Jun 2026 12:30:59 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Anaconda]]></category>
		<category><![CDATA[Jupyter Notebook]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[Transformer]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2500</guid>

					<description><![CDATA[<p>Run a capable large language model entirely on your own machine — private, offline, and with as little as 8 GB of GPU memory. This hands-on guide sets up a clean Python environment, gets CUDA working even on the newest NVIDIA Blackwell cards, loads a 4-bit quantized model from Hugging Face, and builds an interactive chat widget with conversation memory and a live VRAM gauge in JupyterLab. No cloud, no API keys, no data leaving your computer.</p>
<p>The post <a href="https://urbangeoanalytics.com/deploy-local-llm-low-vram-jupyter/">Deploy Your Own Local LLM on Low VRAM in 30 Minutes — A Private Chat Assistant in Jupyter</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-10 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-16 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-15" style="text-align:center;--awb-margin-top:5px;--awb-margin-bottom:5px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-15 hover-type-none"><img decoding="async" width="1693" height="929" title="ILLUS" src="https://urbangeoanalytics.com/wp-content/uploads/2026/06/ILLUS.png" alt class="img-responsive wp-image-2530" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/06/ILLUS-200x110.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/ILLUS-400x219.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/ILLUS-600x329.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/ILLUS-800x439.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/ILLUS-1200x658.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/ILLUS.png 1693w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title"> </div></div></div></div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-17 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-63"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-64" style="--awb-margin-top:-30px;"><ul>
<li>Run a real large language model on your own machine, entirely offline, with as little as 8 GB of GPU memory.</li>
<li>No cloud, no API keys, no data leaving your computer.</li>
<li>Interactive chat widget with conversation memory and a live VRAM gauge</li>
</ul>
</div><div class="fusion-text fusion-text-65 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Cloud chat assistants are convenient, but they come with trade-offs: your queries leave your machine, you depend on someone else&#8217;s uptime and pricing, and the model&#8217;s behaviour can change under you without warning. For research, sensitive data, or simply full control, running a model locally is an appealing alternative. The good news is that modern quantization has made this accessible on modest consumer hardware. A capable 7–8 billion parameter model now fits comfortably on an 8 GB laptop GPU. This tutorial walks through the entire process end to end, using an NVIDIA Blackwell card (RTX 5060, 8 GB) as the worked example — though the approach applies to any recent NVIDIA GPU.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-40 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">1. Setting Up the Environment: Anaconda, a Dedicated Kernel, and the Right CUDA</h2></div><div class="fusion-text fusion-text-66 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Everything starts with a clean, isolated environment. Mixing deep-learning dependencies into your base Python installation is a recipe for version conflicts, so we create a dedicated Conda environment for this project alone. If you have followed our earlier Anaconda setup guide, this will feel familiar.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Open the Anaconda Prompt and create a fresh environment:</p>
</div><div class="fusion-text fusion-text-67"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="bash1" data-enlighter-title="bash">conda create -n localllm python=3.11 -y
conda activate localllm</pre>
</div><div class="fusion-text fusion-text-68 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:15px;--awb-margin-bottom:15px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The single most important, and most overlooked, step is installing the correct build of PyTorch for <em>your specific GPU</em>. This is where most local-LLM attempts fail silently. NVIDIA GPUs each have a &#8220;compute capability&#8221; (an architecture identifier such as sm_86, sm_90, sm_120), and a PyTorch binary only works if it was compiled with kernels for your card&#8217;s architecture. Install the wrong build and you will see CUDA reported as &#8220;available&#8221; while every actual GPU operation crashes — a particularly confusing failure mode.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The newest Blackwell cards (the RTX 50-series, including our RTX 5060) use compute capability sm_120, which older PyTorch wheels do not support. For these cards you need a build compiled against CUDA 12.8 or newer:</p>
</div><div class="fusion-text fusion-text-69"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="bash2" data-enlighter-title="bash">pip install torch --index-url https://download.pytorch.org/whl/cu128</pre>
</div><div class="fusion-text fusion-text-70 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:15px;--awb-margin-bottom:15px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">If you are on an older card (RTX 30- or 40-series), the standard CUDA 12.x wheels are fine. The general rule: match the PyTorch CUDA build to your GPU generation, and when a brand-new card isn&#8217;t yet supported in the stable channel, reach for the nightly build of the matching CUDA version.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Now verify it properly. Do not trust <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">torch.cuda.is_available()</code> alone — it can return <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">True</code> even when no compatible kernels exist. Instead, force an actual computation onto the GPU:</p>
</div><div class="fusion-text fusion-text-71"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="bash3" data-enlighter-title="bash">python -c "import torch; x=torch.randn(1000,1000,device='cuda'); y=x@x;
print('OK', y.device, torch.cuda.get_device_capability(0))"</pre>
</div><div class="fusion-text fusion-text-72 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:15px;--awb-margin-bottom:15px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">A clean <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">OK cuda:0 (12, 0)</code> with no warnings means real GPU compute is working. That is your green light. With the engine confirmed, install the rest of the stack and register the environment as a dedicated Jupyter kernel so the notebook always uses exactly these packages:</p>
</div><div class="fusion-text fusion-text-73"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="bash4" data-enlighter-title="bash">pip install numpy transformers accelerate bitsandbytes jupyterlab ipywidgets ipykernel
python -m ipykernel install --user --name localllm --display-name "Python (localllm)"</pre>
</div><div class="fusion-text fusion-text-74 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:15px;--awb-margin-bottom:15px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Finally, launch JupyterLab <em>from your project directory</em> so your notebook is rooted where you want it rather than in a system folder:</p>
</div><div class="fusion-text fusion-text-75"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="bash5" data-enlighter-title="bash">cd C:\Users\you\Documents\projects
jupyter lab</pre>
</div><div class="fusion-text fusion-text-76 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:15px;--awb-margin-bottom:15px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Open the localhost address provided by jupyter on your navigator and once inside, select the &#8220;Python (localllm)&#8221; kernel. We recommend JupyterLab over the classic Notebook here: it renders interactive widgets reliably out of the box, which matters for the chat interface we build in Section 3.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-16" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-16 hover-type-none"><img decoding="async" width="1970" height="1223" title="localhost kernel" src="https://urbangeoanalytics.com/wp-content/uploads/2026/06/localhost-kernel.png" alt class="img-responsive wp-image-2515" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/06/localhost-kernel-200x124.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/localhost-kernel-400x248.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/localhost-kernel-600x372.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/localhost-kernel-800x497.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/localhost-kernel-1200x745.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/localhost-kernel.png 1970w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">On localhost, choose the kernel we prepared to open a notebook</div></div></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-41 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;"><strong>2. Choosing and Loading the Model: Hugging Face and 4-Bit Quantization</strong></p></h2></div><div class="fusion-text fusion-text-77 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">A model&#8217;s weights have to live in memory, and for modern LLMs they are large. A 7–8 billion parameter model in full 16-bit precision needs roughly 14–16 GB — too much for an 8 GB card. The solution is quantization: storing each weight in 4 bits instead of 16. This shrinks an 8B model to around 5 GB with only a minor quality cost, which is what makes local inference on consumer hardware possible at all.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">We use Hugging Face Transformers together with the bitsandbytes library, which quantizes the model to 4 bits on the fly as it loads. This keeps everything inside your Python kernel — the model object lives in your notebook, you load directly from Hugging Face with optional token authentication, and you can inspect internals if you wish. Hugging Face acts as the model registry: the first load downloads the weights and caches them to disk (under your user folder&#8217;s <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">.cache/huggingface</code>), and every subsequent load reads from that local cache with no network access.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">A note on model choice. There is no single &#8220;best&#8221; small model; it depends on your task and your memory budget. Here is a practical comparison for an 8 GB card:</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3">
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Qwen3 4B Instruct</strong> — the lightweight workhorse. Around 2.7 GB in 4-bit, very fast, strong reasoning and multilingual ability for its size. Ideal as a daily driver for quick questions.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Dolphin 3.0 (Llama 3.1 8B)</strong> — a larger, more capable general-purpose model at around 5–5.5 GB in 4-bit. Built on Llama 3.1 and instruction-tuned by Cognitive Computations, it is designed to put alignment under the user&#8217;s control, making it well suited to research contexts where you define the system prompt and behaviour yourself.</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Other strong candidates</strong> — Phi-4-mini for very light tasks, and Gemma-class models for multilingual writing, depending on what fits your remaining VRAM.</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The rule of thumb: pick the smallest model that does your job well. A 4B model runs noticeably faster than an 8B simply because there are fewer parameters to push through per token, so match model size to task.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The loading code configures 4-bit quantization and reads the model from Hugging Face. We wrap it in a small dropdown so you can switch models without rewriting code:</p>
</div><div class="fusion-text fusion-text-78"><pre class="EnlighterJSRAW" data-enlighter-language="python" data-enlighter-theme="dracula" data-enlighter-group="Python1" data-enlighter-title="Python">import torch, gc
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import ipywidgets as widgets
from IPython.display import display
import warnings
warnings.filterwarnings("ignore", message=".*_check_is_size.*", category=FutureWarning)

MODELS = 

tokenizer = None
model = None

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

dropdown = widgets.Dropdown(options=list(MODELS.keys()), description="Model:",
                            layout=)
load_btn = widgets.Button(description="Load", button_style="primary")
status   = widgets.Output()

def load_model(_=None):
    global tokenizer, model
    model_id = MODELS[dropdown.value]
    with status:
        status.clear_output(); print(f"Loading  …")
    if model is not None:
        del model; model = None
        gc.collect(); torch.cuda.empty_cache()
    tok = AutoTokenizer.from_pretrained(model_id)
    mdl = AutoModelForCausalLM.from_pretrained(
        model_id, quantization_config=bnb_config,
        device_map="cuda:0", dtype=torch.bfloat16,
    )
    mdl.eval()
    tokenizer, model = tok, mdl
    with status:
        print(f"Loaded. VRAM used:  GB")

load_btn.on_click(load_model)
display(widgets.VBox([widgets.HBox([dropdown, load_btn]), status]))</pre>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-17" style="text-align:center;--awb-margin-top:5px;--awb-margin-bottom:5px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-17 hover-type-none"><img decoding="async" width="908" height="170" title="dropdown" src="https://urbangeoanalytics.com/wp-content/uploads/2026/06/dropdown.png" alt class="img-responsive wp-image-2522" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/06/dropdown-200x37.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/dropdown-400x75.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/dropdown-600x112.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/dropdown-800x150.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/dropdown.png 908w" sizes="(max-width: 640px) 100vw, 908px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">The dropdown menu allowing you to choose a model to load</div></div></div></div><div class="fusion-text fusion-text-79 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>If you load a model built on a gated base (such as Llama), you may need to authenticate once with a Hugging Face token via <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">huggingface_hub.login()</code>. Most fine-tuned community models, including the two above, load without one.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-42 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;"><strong>3. Building the Chat Interface: Memory, Context, and a VRAM Gauge</strong></p></h2></div><div class="fusion-text fusion-text-80 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">A loaded model is, by itself, stateless. It has no memory of anything you said previously — each call only sees the text you hand it. To create the experience of a conversation, <em>we</em> must keep the history and re-send it on every turn. Understanding this is the key to using local models well, and it requires distinguishing three concepts that are easy to confuse.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The <strong>context window</strong> is the model&#8217;s hard architectural limit: the maximum number of tokens it can attend to at once, counting both the prompt and the output together. Llama 3.1-based models support up to 128k tokens. The <strong>conversation memory</strong> is not a property of the model at all — it is simply the running list of past turns that we re-inject into the prompt each time, and it consumes part of the context window. The <strong>max new tokens</strong> setting is a cap <em>we choose</em> on how many tokens the model may generate in a single reply; it controls output length only and does not affect what the model can read.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">So the relationship is: the prompt (system message + accumulated history + your new question) plus the reserved output space must all fit inside the context window. The context window is the room; memory is the furniture already in it; max new tokens is the space you set aside for the answer.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">A common misconception is that a larger context window makes the model faster. It is the opposite. A bigger active context costs <em>more</em> VRAM (the key-value cache grows) and runs <em>slower</em>, because each newly generated token must attend over every preceding token. Speed comes from keeping the active context <em>small</em> — short prompts and trimmed history — not large. Reducing max new tokens does not speed up generation either; it simply stops the reply earlier, often mid-thought, since the model does not plan around the limit. The right way to get shorter, faster answers is to instruct the model to be concise via a system prompt, so it produces a complete but brief response.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The widget below puts these ideas into practice. It keeps a <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">chat_history</code> list (the memory), trims it to a fixed number of recent turns (capping context growth), and displays a live VRAM gauge so you can see your headroom and know when to reset. Re-running the cell clears the history — that is your reset.</p>
</div><div class="fusion-text fusion-text-81"><pre class="EnlighterJSRAW" data-enlighter-language="python" data-enlighter-theme="dracula" data-enlighter-group="Python2" data-enlighter-title="Python">import ipywidgets as widgets
from IPython.display import display
import torch

chat_history = []                      # the conversation memory
TOTAL = torch.cuda.get_device_properties(0).total_memory / 1e9
MAX_TURNS = 6                          # cap context: keep last 6 exchanges
SYSTEM = "Be concise. Answer in a few sentences unless asked for detail."

out      = widgets.Output(layout=)
entry    = widgets.Text(placeholder="Type a message…", layout=)
send_btn = widgets.Button(description="Send", button_style="primary")
vram_bar = widgets.FloatProgress(value=0, min=0, max=TOTAL, description="VRAM:")
vram_lbl = widgets.Label()

def refresh_vram():
    used = torch.cuda.memory_allocated() / 1e9
    vram_bar.value = used
    vram_bar.bar_style = ("success" if used < TOTAL*0.6
                          else "warning" if used < TOTAL*0.85 else "danger") vram_lbl.value = f"/ GB ( turns)" def on_send(_=None): global chat_history prompt = entry.value.strip() if not prompt: return if len(chat_history) > MAX_TURNS * 2:          # trim old turns
        chat_history = chat_history[-MAX_TURNS*2:]
    entry.value = ""
    with out:
        print(f"You: ")
    messages = [] + list(chat_history) \
               + []
    text = tokenizer.apply_chat_template(messages, tokenize=False,
                                         add_generation_prompt=True)
    inputs = tokenizer(text, return_tensors="pt").to(model.device)
    with torch.no_grad():
        gen = model.generate(**inputs, max_new_tokens=512,
                             do_sample=False,          # greedy: fast & deterministic
                             pad_token_id=tokenizer.eos_token_id)
    reply = tokenizer.decode(gen[0][inputs["input_ids"].shape[1]:],
                             skip_special_tokens=True)
    chat_history.append()
    chat_history.append()
    with out:
        print(f"Model: \n")
    refresh_vram()

send_btn.on_click(on_send)
refresh_vram()
display(widgets.VBox([out, widgets.HBox([entry, send_btn]),
                      widgets.HBox([vram_bar, vram_lbl])]))</pre>
</div><div class="fusion-text fusion-text-82 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>A few design notes. We use greedy decoding (<code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">do_sample=False</code>) rather than random sampling: it is marginally faster and fully reproducible, with no meaningful quality loss for factual exchanges. The <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">MAX_TURNS</code> value is your direct control over how much the model &#8220;remembers&#8221; versus how lean and fast it stays. And the VRAM gauge turns green, amber, or red as memory fills, giving you a clear signal of when to start a fresh conversation.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-43 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;"><strong>4. The Assistant in Action</strong></p></h2></div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-18" style="text-align:center;--awb-margin-top:5px;--awb-margin-bottom:5px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-18 hover-type-none"><img decoding="async" width="1853" height="626" title="assistant1" src="https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant1.png" alt class="img-responsive wp-image-2525" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant1-200x68.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant1-400x135.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant1-600x203.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant1-800x270.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant1-1200x405.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant1.png 1853w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">The loaded assistant with VRAM use and a reset function</div></div></div></div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-19" style="text-align:center;--awb-margin-top:5px;--awb-margin-bottom:5px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-19 hover-type-none"><img decoding="async" width="1842" height="624" title="assistant2" src="https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant2.png" alt class="img-responsive wp-image-2526" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant2-200x68.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant2-400x136.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant2-600x203.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant2-800x271.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant2-1200x407.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant2.png 1842w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">Let's try it with a question and then try the memory</div></div></div></div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-20" style="text-align:center;--awb-margin-top:5px;--awb-margin-bottom:5px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-20 hover-type-none"><img decoding="async" width="1849" height="624" title="assistant3" src="https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant3.png" alt class="img-responsive wp-image-2527" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant3-200x67.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant3-400x135.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant3-600x202.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant3-800x270.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant3-1200x405.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/06/assistant3.png 1849w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">Everything works well including the memory, well done!</div></div></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-44 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;"><strong>Conclusion: Why Local Matters — and What Comes Next</strong></p></h2></div><div class="fusion-text fusion-text-83 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">What we have built is small but genuinely yours. Every conversation lives only in your computer&#8217;s memory, inside the running notebook kernel. Nothing is written to disk, nothing is sent anywhere, and nothing is logged. Close the kernel and the entire conversation simply vanishes — the only thing that persists is the downloaded model weights in your local cache. For sensitive research data, confidential analysis, or simply peace of mind, this is a meaningful difference from any cloud service.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Beyond privacy, running locally brings other advantages. You are not subject to per-token billing or rate limits, so you can experiment freely. You are insulated from silent model changes and deprecations — your model behaves the same tomorrow as it does today. And with community fine-tunes such as Dolphin, you control the system prompt and the model&#8217;s alignment yourself, rather than inheriting a one-size-fits-all policy. With fewer built-in guardrails, these models will engage with a wider range of legitimate research and technical questions, which can be valuable in specialist domains where general-purpose assistants are overly cautious — a freedom that naturally comes with the responsibility to use it sensibly.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">This is only the foundation. In future posts we will extend this local assistant in several directions. We will give it <strong>web browsing</strong>, so it can retrieve current information rather than relying solely on its training. We will explore an <strong>expert mode</strong>, pre-loading the context with domain knowledge — for instance a corpus of spatial-analysis references — so the assistant answers as a specialist in your field. And we will look at <strong>containerizing</strong> the whole setup with Docker so it can be deployed on a dedicated GPU server or in the cloud, turning this notebook prototype into a private assistant you can embed directly in your own website.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">For now, you have a capable, private language model running on hardware you already own, set up in about half an hour. Learn it, build on it, and apply it to your own work.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-18 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-84"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--7" data-awb-toc-id="7" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-21 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/deploy-local-llm-low-vram-jupyter/">Deploy Your Own Local LLM on Low VRAM in 30 Minutes — A Private Chat Assistant in Jupyter</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/deploy-local-llm-low-vram-jupyter/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>SAGAI v2.0 — A Unified Multi-Model Notebook for Streetscape Analysis</title>
		<link>https://urbangeoanalytics.com/sagai-v2-multi-model-streetscape-analysis-uvlm/</link>
					<comments>https://urbangeoanalytics.com/sagai-v2-multi-model-streetscape-analysis-uvlm/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Thu, 21 May 2026 10:11:18 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[GIS]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Llava]]></category>
		<category><![CDATA[Qwen]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2483</guid>

					<description><![CDATA[<p>SAGAI v2.0 consolidates the full streetscape analysis pipeline into a single Google Colab notebook and replaces the inline LLaVA-only inference code with the UVLM package, enabling multi-model benchmarking across 11 VLM checkpoints. New features include a multi-task prompt builder, consensus validation with majority voting, chain-of-thought reasoning, truncation detection, interactive Folium maps, view-direction filtering, and support for loading existing polygons as study area boundaries.</p>
<p>The post <a href="https://urbangeoanalytics.com/sagai-v2-multi-model-streetscape-analysis-uvlm/">SAGAI v2.0 — A Unified Multi-Model Notebook for Streetscape Analysis</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-11 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-19 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-22" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-22 hover-type-none"><img decoding="async" width="1760" height="545" title="e4e3b0b4-83a7-4933-ba0b-ef1775beacc6" src="https://urbangeoanalytics.com/wp-content/uploads/2026/05/e4e3b0b4-83a7-4933-ba0b-ef1775beacc6.png" alt class="img-responsive wp-image-2489" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/05/e4e3b0b4-83a7-4933-ba0b-ef1775beacc6-200x62.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/e4e3b0b4-83a7-4933-ba0b-ef1775beacc6-400x124.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/e4e3b0b4-83a7-4933-ba0b-ef1775beacc6-600x186.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/e4e3b0b4-83a7-4933-ba0b-ef1775beacc6-800x248.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/e4e3b0b4-83a7-4933-ba0b-ef1775beacc6-1200x372.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/e4e3b0b4-83a7-4933-ba0b-ef1775beacc6.png 1760w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title"> </div></div></div></div><div class="fusion-text fusion-text-85"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-86" style="--awb-margin-top:-30px;"><ul>
<li>SAGAI v2.0 merges the previous four-module notebook architecture into a <strong>single unified Google Colab notebook</strong> (SAGAI.ipynb) organized in six sequential blocks.</li>
<li>The inline LLaVA-only inference code is replaced by the <strong>UVLM package</strong> (Universal Vision-Language Model Loader), installed automatically from GitHub, providing access to <strong>11 VLM checkpoints</strong> across two model families.</li>
<li>New capabilities include a <strong>multi-task prompt builder</strong>, <strong>consensus validation</strong> with majority voting, <strong>chain-of-thought reasoning</strong>, <strong>truncation detection</strong>, <strong>interactive Folium maps</strong>, <strong>view-direction filtering</strong>, and support for <strong>loading an existing study area polygon</strong>.</li>
</ul>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-45 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Introduction</h2></div><div class="fusion-text fusion-text-87 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">SAGAI (Streetscape Analysis with Generative Artificial Intelligence) is an open-source workflow for scoring and mapping street-level urban environments using vision-language models and open geospatial data. Since its initial release, SAGAI has been structured as a set of independent Colab notebooks, one per pipeline stage, each relying on its own dependencies and documentation.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">SAGAI v2.0 is a major release that consolidates the entire pipeline into a single notebook and replaces the custom inference code with the UVLM package. Where previous versions were tied to a single LLaVA checkpoint with handwritten inference logic, SAGAI v2.0 delegates all vision-language model loading, prompting, and evaluation to UVLM&#8217;s unified interface. This makes the scoring engine model-agnostic: users can select from 11 VLM checkpoints spanning the LLaVA-NeXT and Qwen2.5-VL families, compare their performance on identical tasks, and benefit from features such as consensus validation, reasoning traces, and truncation diagnostics; all within the same notebook.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Beyond the inference engine, v2.0 introduces structural and functional changes across the entire pipeline: a unified six-block architecture, interactive HTML mapping via Folium, view-direction filtering for aggregation, and the ability to load an existing polygon as a study area boundary instead of defining a bounding box manually.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">This post details the architectural changes, the UVLM integration, and the new features introduced in SAGAI v2.0.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-46 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">1. From Four Notebooks to One: The Unified Architecture</h2></div><div class="fusion-text fusion-text-88 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Previous SAGAI releases were organized as four independent Colab notebooks — one for street sampling, one for image retrieval, one for VLM inference, and one for aggregation and mapping — each accompanied by a separate NOTICE file documenting its dependencies and usage. This modular design was useful for development but introduced friction in practice: users had to manage file paths between notebooks, track four separate environments, and consult multiple documentation files.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">SAGAI v2.0 merges all four stages into a single notebook (SAGAI.ipynb) structured as six sequential blocks. The pipeline flows from study area definition through street sampling, image downloading, VLM scoring, and mapping, with all intermediate data passed directly between blocks in the same runtime session. The separate per-module NOTICE files and the standalone requirements file (requirements_sagai_module_3_v1-0.txt) have been removed — dependency management is now handled automatically by the UVLM package installation.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-23" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-23 hover-type-none"><img decoding="async" width="2000" height="948" title="pipeline details" src="https://urbangeoanalytics.com/wp-content/uploads/2026/05/pipeline-details-scaled.png" alt class="img-responsive wp-image-2480" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/05/pipeline-details-300x142.png 300w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/pipeline-details-768x364.png 768w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/pipeline-details-1024x486.png 1024w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/pipeline-details-1536x728.png 1536w, https://urbangeoanalytics.com/wp-content/uploads/2026/05/pipeline-details-scaled.png 2000w" sizes="(max-width: 2000px) 100vw, 2000px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">Diagram of the six-block architecture</div></div></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-47 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">2. Study Area Definition: Bounding Box or Existing Polygon</h2></div><div class="fusion-text fusion-text-89 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">In previous versions, the study area was defined exclusively by a bounding box in WGS84 coordinates. SAGAI v2.0 retains this option but adds the ability to draw your own polygon or to load an existing polygon; for example, a GeoPackage representing a neighborhood, municipality, or custom boundary. When a polygon is provided, the street sampling step extracts the OpenStreetMap network within that geometry rather than a rectangular extent. This makes it straightforward to work with irregular administrative boundaries or user-defined study zones without manually computing bounding coordinates.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-48 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">3. UVLM Integration: From Single-Model Inference to Multi-Model Benchmarking</h2></div><div class="fusion-text fusion-text-90 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The most significant change in SAGAI v2.0 is the replacement of the inline inference code with the <a class="keychainify-checked" href="https://github.com/perezjoan/UVLM/tree/main">UVLM package</a>. In previous versions, Blocks 3 through 5 contained custom code for loading a single LLaVA checkpoint, constructing prompts, running inference, and parsing outputs. This logic was tightly coupled to one model architecture and required manual maintenance when Hugging Face APIs or model formats changed.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">SAGAI v2.0 installs UVLM directly from its GitHub repository at the start of the notebook. All model loading, prompt formatting, inference execution, response parsing, and batch processing are delegated to UVLM&#8217;s API. The inline inference code has been entirely removed.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Through UVLM, SAGAI v2.0 supports 11 VLM checkpoints across two model families:</p>
<ul class="&#091;li_&amp;&#093;:mb-0 &#091;li_&amp;&#093;:mt-1 &#091;li_&amp;&#093;:gap-1 &#091;&amp;:not(:last-child)_ul&#093;:pb-1 &#091;&amp;:not(:last-child)_ol&#093;:pb-1 list-disc flex flex-col gap-1 pl-8 mb-3">
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>LLaVA-NeXT</strong> — Mistral 7B, Vicuna 7B, Vicuna 13B, 34B, LLaMA3 8B, 72B, 110B</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Qwen2.5-VL</strong> — 3B Instruct, 7B Instruct, 32B Instruct, 72B Instruct</li>
</ul>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">UVLM&#8217;s dual-backend abstraction automatically detects the model family and routes inference to the correct pipeline — LlavaNextProcessor for LLaVA models, AutoProcessor with process_vision_info for Qwen models — so users switch between architectures by changing a single model selection, with no modification to the rest of the notebook.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Quantization is handled through UVLM&#8217;s built-in support for 4-bit, 8-bit, and FP16 precision via BitsAndBytes. Models up to 34B parameters can run on a single Colab GPU (T4 or A100) with 4-bit quantization.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-49 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">4. Multi-Task Prompt Builder</h2></div><div class="fusion-text fusion-text-91 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">UVLM provides a widget-based prompt builder that SAGAI v2.0 exposes directly in the notebook. Users can define up to 10 analysis tasks per run, each with its own prompt, response type (numeric, category, boolean, or text), and label. This replaces the previous approach of selecting from a small set of hardcoded tasks (T1, T2, T3) or manually editing prompt strings in the code.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Tasks are configured interactively before execution and applied uniformly across all images in the batch. Each task produces its own column in the output CSV file.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-24" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-24 hover-type-none"><img decoding="async" width="866" height="1063" title="image2" src="https://urbangeoanalytics.com/wp-content/uploads/2026/03/image2.png" alt class="img-responsive wp-image-2320" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/03/image2-200x245.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image2-400x491.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image2-600x736.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image2-800x982.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image2.png 866w" sizes="(max-width: 640px) 100vw, 866px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">UVLM prompt builder</div></div></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-50 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">5. Consensus Validation</h2></div><div class="fusion-text fusion-text-92 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">SAGAI v2.0 inherits UVLM&#8217;s consensus validation mechanism. Each analysis task can be run 2 to 5 times per image, and the final score is determined by majority voting across the repeated inferences. NA values from failed parses are filtered before voting. An agreement ratio is recorded alongside the final score, providing a built-in measure of prediction reliability without any external validation step.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-51 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">6. Chain-of-Thought Reasoning and Truncation Detection</h2></div><div class="fusion-text fusion-text-93 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">UVLM supports two approaches to chain-of-thought (CoT) reasoning, both available in SAGAI v2.0. Users can write task prompts that explicitly request step-by-step reasoning and adjust the token budget (up to 1,500 tokens) to allow the model sufficient generation space. Alternatively, a built-in CoT reference mode can be enabled per task, which triggers a standardized reasoning template with a fixed 1,024-token budget. In both cases, the reasoning trace is stored in a dedicated column in the output CSV for inspection.</p>
<p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Truncation detection is performed automatically after every inference call. The exact number of generated tokens is compared against the token limit, and truncated responses are flagged in per-task CSV columns. This allows users to identify tasks where the token budget is insufficient without post-hoc analysis.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-52 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">7. Interactive Mapping with Folium</h2></div><div class="fusion-text fusion-text-94 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Previous SAGAI versions generated static thematic maps using Matplotlib. SAGAI v2.0 replaces these with interactive HTML maps built with Folium. Point-level and street-segment-level scores are rendered as interactive layers that can be panned, zoomed, and queried directly in the browser. This is particularly useful for exploratory analysis and for sharing results with collaborators who do not use GIS software.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-53 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">8. View-Direction Filtering for Aggregation</h2></div><div class="fusion-text fusion-text-95 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">Google Street View images are typically downloaded in multiple compass directions at each sampling point (e.g., front, back, left, right). In previous versions, all views were aggregated together when computing point- or street-level scores. SAGAI v2.0 introduces a view filter that allows users to select which directions to include in the aggregation — for example, scoring only left-side and right-side views to focus on building facades, or only front views to capture the pedestrian perspective along the street axis. This filter is applied at the aggregation stage and does not affect the scoring step itself.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-54 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">9. Resume-Safe Batch Processing</h2></div><div class="fusion-text fusion-text-96 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal leading-&#091;1.7&#093;">The batch execution engine inherited from UVLM provides resume-safe processing with checkpoint saving every 3 images. If a Colab session is interrupted — due to a timeout, a runtime reset, or a connectivity issue — the notebook can be re-executed and will automatically skip already-processed images. New tasks added between runs trigger automatic CSV schema upgrading, so the output file grows incrementally without losing previous results.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-55 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">10. References and Links</h2></div><div class="fusion-text fusion-text-97 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><ul>
<li class="font-claude-response-body whitespace-normal break-words pl-2">SAGAI v2.0 on GitHub: <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://github.com/perezjoan/SAGAI">https://github.com/perezjoan/SAGAI</a></li>
<li class="font-claude-response-body whitespace-normal break-words pl-2">UVLM on GitHub: <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://github.com/perezjoan/UVLM">https://github.com/perezjoan/UVLM</a></li>
<li class="font-claude-response-body whitespace-normal break-words pl-2">Perez, J. and Fusco, G. (2025). <em>Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes.</em> Geomatica, 77(2), 100063. <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://www.sciencedirect.com/science/article/pii/S1195103625000199">https://www.sciencedirect.com/science/article/pii/S1195103625000199</a></li>
<li class="font-claude-response-body whitespace-normal break-words pl-2">Perez, J. and Fusco, G. (2026). <em>UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking.</em> arXiv:2603.13893. <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://arxiv.org/abs/2603.13893">https://arxiv.org/abs/2603.13893</a></li>
</ul>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-20 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-98"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--8" data-awb-toc-id="8" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-25 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/sagai-v2-multi-model-streetscape-analysis-uvlm/">SAGAI v2.0 — A Unified Multi-Model Notebook for Streetscape Analysis</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/sagai-v2-multi-model-streetscape-analysis-uvlm/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>UVLM v3.0.0: From Colab Notebook to Python Package — Run Vision-Language Models Anywhere</title>
		<link>https://urbangeoanalytics.com/uvlm-python-package-vision-language-models/</link>
					<comments>https://urbangeoanalytics.com/uvlm-python-package-vision-language-models/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Thu, 23 Apr 2026 07:25:41 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Package]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Google Colab]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Jupyter Notebook]]></category>
		<category><![CDATA[Llava]]></category>
		<category><![CDATA[Qwen]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2442</guid>

					<description><![CDATA[<p>UVLM v3.0.0 turns a Colab notebook into a full Python package. Run vision-language models locally, in notebooks, or scripts with a simple API and no setup complexity.</p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-python-package-vision-language-models/">UVLM v3.0.0: From Colab Notebook to Python Package — Run Vision-Language Models Anywhere</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-12 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-21 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-26" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-26 hover-type-none"><img decoding="async" width="1619" height="971" title="flag fig" src="https://urbangeoanalytics.com/wp-content/uploads/2026/04/flag-fig.png" alt class="img-responsive wp-image-2469" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/04/flag-fig-200x120.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/flag-fig-400x240.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/flag-fig-600x360.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/flag-fig-800x480.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/flag-fig-1200x720.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/flag-fig.png 1619w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title"> </div></div></div></div><div class="fusion-text fusion-text-99"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-100" style="--awb-margin-top:-30px;"><ul>
<li><strong data-start="64" data-end="88">UVLM is now a pip-installable Python package </strong>— no longer tied to Google Colab</li>
<li><strong data-start="64" data-end="88">Run on your own GPU </strong>with a local Jupyter notebook, or keep using Colab for free</li>
<li><strong data-start="64" data-end="88">Same tool, more flexibility </strong>— three lines of Python to load a model and analyse images</li>
</ul>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-text fusion-text-101 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>When we released UVLM in March 2026, it was a Google Colab notebook. You opened it in your browser, picked a model, typed your prompts, and ran your images — all without installing anything. That simplicity was the point: a tool that anyone could use to load and compare Vision-Language Models, regardless of their technical setup.</p>
<p>But we kept hearing the same requests. Can I run this on my own machine? Can I call UVLM from a script? Can I integrate it into an existing pipeline? The answer was always the same: not easily. The entire tool lived inside a single notebook, with all the logic packed into three massive code cells. Moving it anywhere else meant copy-pasting thousands of lines and untangling global variables.</p>
<p>Version 3.0.0 changes that. UVLM is now a proper Python package.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-56 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">What Changed</h2></div><div class="fusion-text fusion-text-102 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>The core logic — model loading, dual-backend inference, response parsing, consensus validation, batch processing — has been extracted from the notebook into eight standalone Python modules. These modules have no dependency on Google Colab, no global variables, and no widget code. They are plain Python functions that accept arguments and return results.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-27" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-27 hover-type-none"><img decoding="async" width="2000" height="1162" title="UVLM package blogpost figure 1" src="https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-package-blogpost-figure-1-scaled.png" alt class="img-responsive wp-image-2444" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-package-blogpost-figure-1-200x116.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-package-blogpost-figure-1-400x232.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-package-blogpost-figure-1-600x349.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-package-blogpost-figure-1-800x465.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-package-blogpost-figure-1-1200x697.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-package-blogpost-figure-1-scaled.png 2000w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title"> </div></div></div></div><div class="fusion-text fusion-text-103 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>The package is installed from GitHub in one line:</p>
</div><div class="fusion-text fusion-text-104 fusion-text-no-margin" style="--awb-margin-top:1px;--awb-margin-bottom:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="python" data-enlighter-theme="dracula" data-enlighter-group="Python1" data-enlighter-title="Python">pip install git+https://github.com/perezjoan/UVLM.git</pre>
</div><div class="fusion-text fusion-text-105 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:5px;--awb-margin-bottom:25px;"><p>On Google Colab, this happens automatically in the first cell of the Colab notebook. On your local machine, you run it once in a terminal and you are done.</p>
<p>Nothing changed in how UVLM analyses images. The same 11 model checkpoints are supported (LLaVA-NeXT and Qwen2.5-VL, from 3B to 110B parameters). The same parsing logic, the same consensus validation, the same truncation detection. If you had a workflow built on v2.2.2, the outputs will be identical.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-57 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Three Ways to Use UVLM</h2></div><div class="fusion-text fusion-text-106 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p><strong>Google Colab — Zero Install</strong></p>
<p>This is the same experience as before. Open the Colab notebook, select a GPU runtime, and start working. The notebook installs the UVLM package automatically. Images are loaded from Google Drive. Nothing has changed for Colab users, except that the code running behind the widgets is now cleaner and easier to maintain.</p>
<p><strong>Local Jupyter Notebook — Your GPU, Your Data</strong></p>
<p>If you have an NVIDIA GPU on your workstation (or access to a GPU server), you can now run UVLM locally. The local Jupyter notebook provides the same widget-based interface — model selection dropdown, prompt builder form, batch execution button — but images are read from your local filesystem and results are saved locally. No Google account needed, no data leaves your machine.</p>
<p>This matters for researchers working with sensitive imagery (medical, security, proprietary datasets) or for anyone who wants faster and more reliable model loading than what Colab&#8217;s network provides.</p>
<p><strong>Python Script — Full Programmatic Control</strong></p>
<p>For integration into larger pipelines, UVLM now exposes a clean API. Three lines of code replace the entire notebook workflow:</p>
</div><div class="fusion-text fusion-text-107 fusion-text-no-margin" style="--awb-margin-top:1px;--awb-margin-bottom:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="python" data-enlighter-theme="dracula" data-enlighter-group="Python2" data-enlighter-title="Python">from uvlm import load_model, run_inference, parse_response
ctx = load_model("[Qwen] Qwen2.5-VL 7B Instruct", precision="4bit")
raw, tokens = run_inference("photo.jpg", "Count the cars", ctx)
result = parse_response(raw, "numeric")</pre>
</div><div class="fusion-text fusion-text-108 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:5px;--awb-margin-bottom:25px;"><p>The `load_model()` function returns a context dictionary containing the model, processor, backend type, and device information. This dictionary is passed to every subsequent function — no global state, no hidden side effects. You can load multiple models in the same session and switch between them by passing different context objects.</p>
<p>For batch processing, `run_batch()` handles the full pipeline:</p>
</div><div class="fusion-text fusion-text-109 fusion-text-no-margin" style="--awb-margin-top:1px;--awb-margin-bottom:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="python" data-enlighter-theme="dracula" data-enlighter-group="Python3" data-enlighter-title="Python">from uvlm import load_model
from uvlm.batch import run_batch

ctx = load_model("[Qwen]  Qwen2.5-VL 7B Instruct", precision="4bit")
df = run_batch(
    model_ctx=ctx,
    task_specs=my_tasks,
    image_folder="./images",
    output_path="./results.csv",
)
</pre>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-28" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-28 hover-type-none"><img decoding="async" width="2000" height="926" title="UVLM deploy blogpost figure 2" src="https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-deploy-blogpost-figure-2-scaled.png" alt class="img-responsive wp-image-2457" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-deploy-blogpost-figure-2-200x93.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-deploy-blogpost-figure-2-400x185.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-deploy-blogpost-figure-2-600x278.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-deploy-blogpost-figure-2-800x370.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-deploy-blogpost-figure-2-1200x556.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/04/UVLM-deploy-blogpost-figure-2-scaled.png 2000w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title"> </div><p class="awb-imageframe-caption-text"> </p></div></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-58 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Under the Hood: Package Structure</h2></div><div class="fusion-text fusion-text-110 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>The monolithic notebook has been split into eight modules, each with a single responsibility:</p>
<p><em>registry.py</em> holds the model dictionary — 11 checkpoints with their backend type and <strong>HuggingFace checkpoint ID</strong>. Adding a new model is one line in a dictionary.</p>
<p><em>loader.py</em> contains the `load_model()` function. It handles quantisation configuration (4-bit, 8-bit, FP16), device placement (single GPU, auto, CPU offload), and the LLaVA vs Qwen branching logic. It returns a dictionary — not a set of global variables.</p>
<p><em>inference.py</em> contains `run_inference()`, the dual-backend forward pass. It accepts a model context dictionary and returns the raw response plus the exact token count as a tuple. The full LLaVA response cleaning logic and the full Qwen token-trimming pipeline are preserved exactly as they were.</p>
<p><em>parsers.py</em> holds the four response parsers (numeric, category, boolean, text) and the advanced reasoning parser. These are pure functions with zero dependencies beyond Python&#8217;s standard library.</p>
<p><em>consensus.py</em> contains the majority voting logic. <em>batch.py</em> handles folder iteration, CSV writing, resume mode, and schema upgrading. <em>prompts.py</em> stores the task type definitions and the chain-of-thought templates. <em>utils.py</em> provides seed management, environment detection, and <strong>HuggingFace token</strong> retrieval.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-59 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Getting Started</h2></div><div class="fusion-text fusion-text-111 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p><strong>On Colab</strong>: Open the notebook from GitHub and run the three blocks as before. The package installs itself.</p>
<p><strong>Locally</strong>: First, install PyTorch with CUDA support matching your GPU driver (check with `nvidia-smi`). For example, with CUDA 12.8+:</p>
</div><div class="fusion-text fusion-text-112 fusion-text-no-margin" style="--awb-margin-top:1px;--awb-margin-bottom:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="python" data-enlighter-theme="dracula" data-enlighter-group="Python4" data-enlighter-title="Python">pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install git+https://github.com/perezjoan/UVLM.git
</pre>
</div><div class="fusion-text fusion-text-113 fusion-text-no-margin" style="--awb-margin-top:1px;--awb-margin-bottom:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="python" data-enlighter-theme="dracula" data-enlighter-group="Python4" data-enlighter-title="Python">pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128
pip install git+https://github.com/perezjoan/UVLM.git
</pre>
</div><div class="fusion-text fusion-text-114 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:5px;--awb-margin-bottom:25px;"><p>Then open the local Jupyter notebook.</p>
<p>You get the same dropdown menus, the same prompt builder form, the same batch execution. The only difference is that you type a local path for your image folder instead of a Google Drive path.</p>
<p>For HuggingFace authentication (needed for some gated models like LLaMA3-based checkpoints), either set the `HF_TOKEN` environment variable or run `huggingface-cli login` once in your terminal.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-60 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">What Is Next</h2></div><div class="fusion-text fusion-text-115 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>The package architecture makes it much easier to add new VLM families. InternVL, BLIP-2, CogVLM, DeepSeek-VL, and Molmo are planned for future releases — each one requires implementing the backend-specific sections of the inference function and adding entries to the registry, without touching the rest of the codebase.</p>
<p>We are also working on multi-GPU batching for parallel inference across images, video frame analysis support, and integration with the SAGAI workflow for automated streetscape analysis.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-61 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Links</h2></div><div class="fusion-text fusion-text-116 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Source code: <a class="keychainify-checked" href="https://github.com/perezjoan/UVLM">github.com/perezjoan/UVLM</a></p>
<p>Paper: <a class="keychainify-checked" href="https://arxiv.org/abs/2603.13893">arXiv preprint</a> — Perez &amp; Fusco (2026)</p>
<p>UVLM page on this site: urbangeoanalytics.com › Software &amp; Algorithms › <a class="keychainify-checked" href="https://urbangeoanalytics.com/algorithms-softwares/uvlm-universal-vision-language-model-loader/">UVLM</a></p>
<p>Previous blog post: <a class="keychainify-checked" href="https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/">Introducing UVLM: A Free Tool to Compare AI Models That Understand Images</a></p>
</div><div class="fusion-title title fusion-title-62 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Citation</h2></div><div class="fusion-text fusion-text-117 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>If you use UVLM in your work, please cite:</p>
<p>Perez, J. &amp; Fusco, G. (2026). <em>UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking.</em> arXiv:2603.13893</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-22 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-118"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--9" data-awb-toc-id="9" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-29 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-python-package-vision-language-models/">UVLM v3.0.0: From Colab Notebook to Python Package — Run Vision-Language Models Anywhere</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/uvlm-python-package-vision-language-models/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Introducing UVLM: A Free Tool to Compare AI Models That Understand Images</title>
		<link>https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/</link>
					<comments>https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Tue, 17 Mar 2026 14:23:58 +0000</pubDate>
				<category><![CDATA[Intermediate]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[Benchmarking]]></category>
		<category><![CDATA[Chain-of-Thought]]></category>
		<category><![CDATA[Google Colab]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Llava]]></category>
		<category><![CDATA[Multimodal AI]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[Qwen]]></category>
		<category><![CDATA[UVLM]]></category>
		<category><![CDATA[VLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2356</guid>

					<description><![CDATA[<p>UVLM is a free, open-source tool for loading, testing, and comparing Vision-Language Models on custom image analysis tasks. Running entirely in Google Colab, it lets researchers and practitioners benchmark multiple AI models using the same prompts and images — no coding, no GPU ownership, no model-specific pipelines. This post explains what VLMs are, why comparing them matters, and how to get started in five minutes.</p>
<p>The post <a href="https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/">Introducing UVLM: A Free Tool to Compare AI Models That Understand Images</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-13 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-23 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-30" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-30 hover-type-none"><img decoding="async" width="1536" height="595" title="uvlm" src="https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm.png" alt class="img-responsive wp-image-2342" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-200x77.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-400x155.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-600x232.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-800x310.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-1200x465.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm.png 1536w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">uvlm</div></div></div></div><div class="fusion-text fusion-text-119"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-120" style="--awb-margin-top:-30px;"><ul>
<li><strong>New open-source release: UVLM v2.2.2</strong> — compare Vision-Language Models from a single notebook</li>
<li><strong>11 AI models</strong>, 5 analysis tasks, 120 test images — all benchmarked with one tool</li>
<li><strong>No coding, no installation</strong> — runs in Google Colab with a free account</li>
</ul>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-text fusion-text-121 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Imagine you have thousands of street photographs and you need to answer the same questions about each one: how many cars are parked? Is there a sidewalk? How long is the building frontage? Hiring someone to go through every image manually would take weeks. Training a custom computer vision model would take months. But what if you could simply ask an AI model these questions in plain English — and get structured, usable answers back?</p>
<p>That is exactly what Vision-Language Models do. And today, we are releasing UVLM — an open-source tool that makes it easy to load, test, and compare these models, all from a single notebook in your browser.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-63 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">What Are Vision-Language Models?</h2></div><div class="fusion-text fusion-text-122 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Vision-Language Models (VLMs) are AI systems that can look at an image and answer questions about it in natural language. Unlike traditional computer vision, which requires training a separate model for every task (one for counting cars, another for detecting sidewalks, a third for classifying buildings), a VLM handles all of these through text prompts. You write a question, attach a photo, and the model responds.</p>
<p>For example, you can ask a VLM: “Count all motor vehicles visible in this image” and it will answer “3”. You can ask the same model “Is there a sidewalk along the street frontage?” and it will answer “yes”. You can even ask it to estimate the length of a building facade in meters — a task that requires the model to identify reference objects (like parked cars), estimate their size, and reason about perspective. All of this from a single model, with no retraining and no labelled dataset.</p>
<p>The catch is that there are many VLM families available (LLaVA, Qwen, InternVL, BLIP-2, and more), and each one works differently under the hood. They use different image encoders, different tokenisation strategies, and different code to run. If you want to know which model is best for your specific task, you normally have to write separate code for each one — a tedious and error-prone process.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-64 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">This Is the Problem UVLM Solves</h2></div><div class="fusion-text fusion-text-123 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>UVLM (Universal Vision-Language Model Loader) is a free, open-source tool that lets you load, configure, and compare multiple VLM architectures using the same prompts and the same evaluation protocol — without writing any model-specific code. It runs entirely in Google Colab, which means you do not need to install anything on your computer or own a GPU. A free Google account is all you need.</p>
<p>The idea is simple: you pick a model from a dropdown menu, type your analysis questions into a form, point the tool at a folder of images, and hit run. UVLM handles all the technical details — the processor classes, the tokenisation, the generation settings, the output parsing — and delivers a clean CSV file with one row per image and one column per task. If you want to try a different model, you just switch the dropdown and run again. Same prompts, same images, same output format. Now you can compare.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-31" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-31 hover-type-none"><img decoding="async" width="1190" height="823" title="image1" src="https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1.png" alt class="img-responsive wp-image-2319" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-200x138.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-400x277.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-600x415.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-800x553.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1.png 1190w" sizes="(max-width: 640px) 100vw, 1190px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">The 3 blocks structure of UVLM Loader</div></div></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-65 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">A Practical Example: Scoring 120 Street Photographs</h2></div><div class="fusion-text fusion-text-124 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>To demonstrate what UVLM can do, we benchmarked 8 different models on 120 street-level photographs of French urban frontages. Each image was analysed on five tasks: counting vehicles, detecting sidewalks, counting pedestrian entrances, estimating the street frontage length in meters, and classifying the vegetation type. That is 16 model configurations (each model tested in standard and advanced reasoning modes), 120 images, and 5 tasks per image — all processed and compared through UVLM.</p>
<p>The results were revealing. The largest model (LLaVA 34B, with 34 billion parameters) actually ranked last overall. A much smaller model (LLaVA Vicuna 7B) outperformed it significantly and ran on a free Google Colab GPU. The best overall results came from Qwen 32B with chain-of-thought reasoning enabled, which achieved 88% proximity to human expert annotations across all five tasks. Without UVLM, discovering these differences would have required writing and debugging eight separate inference pipelines.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-66 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Who Is UVLM For?</h2></div><div class="fusion-text fusion-text-125 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>UVLM was designed for anyone who works with images and wants to extract structured information from them at scale — without becoming a machine learning engineer. If you are an urban planner evaluating streetscape quality across a city, UVLM lets you score thousands of street photographs using natural language prompts. If you are an environmental researcher classifying vegetation from field photographs, UVLM lets you test which AI model gives the most reliable results for your specific classification scheme. If you are an infrastructure inspector processing damage assessment photographs, UVLM lets you set up automated counting and scoring tasks and run them across your entire image archive.</p>
<p>The tool is also valuable for AI researchers who need a controlled benchmarking environment. Because UVLM ensures that every model receives exactly the same prompt and is evaluated with the same metrics, it produces fair, reproducible comparisons. The consensus validation feature (running each task multiple times and taking a majority vote) addresses the inherent randomness of AI outputs, and the truncation detection feature flags when a model’s response was cut off before it could finish — a common but often invisible source of errors.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-67 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">How to Get Started</h2></div><div class="fusion-text fusion-text-126 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Getting started takes about five minutes. Open the UVLM notebook from GitHub (the link is below), connect to a GPU runtime in Google Colab, and run the first block to load a model. The second block gives you a form where you type your analysis questions — no coding required. The third block processes your images and saves the results as a CSV file on your Google Drive.</p>
<p>The tool currently supports 11 model checkpoints from two major families (LLaVA-NeXT and Qwen2.5-VL), ranging from 3 billion to 110 billion parameters. Models up to 34B can run on a single free-tier Colab GPU with 4-bit quantisation. Advanced features include consensus validation (2–5 runs per task with majority voting), chain-of-thought reasoning for complex tasks, and automatic truncation detection.</p>
<p>UVLM is released under the Apache 2.0 open-source licence. You can use it, modify it, and build on it for any purpose — academic or commercial.</p>
</div><div class="fusion-text fusion-text-127 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2>Links</h2>
<p><strong>Source code: </strong><a class="keychainify-checked" href="https://github.com/perezjoan/UVLM">github.com/perezjoan/UVLM</a></p>
<p><strong>Paper: </strong><a class="keychainify-checked" href="https://arxiv.org/abs/2603.13893">arXiv preprint — Perez &amp; Fusco (2026)</a></p>
<p><strong>UVLM page on this site: </strong><a class="keychainify-checked" href="https://urbangeoanalytics.com/algorithms-softwares/uvlm-universal-vision-language-model-loader/">urbangeoanalytics.com › Softwares &amp; Algorithms › UVLM</a></p>
<p><strong>Benchmark dataset: </strong><a class="keychainify-checked" href="https://zenodo.org/records/18959690">Zenodo — 120 street-view images</a></p>
<h2>Citation</h2>
<p>If you use UVLM in your work, please cite:</p>
<p><em>Perez, J. &amp; Fusco, G. (2026). UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking. arXiv:2603.13893</em></p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-24 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-128"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--10" data-awb-toc-id="10" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-32 hover-type-zoomout"><img decoding="async" width="1536" height="1024" title="blog lvl2" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15.png" alt class="img-responsive wp-image-1687" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/">Introducing UVLM: A Free Tool to Compare AI Models That Understand Images</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
