<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>Open Source Archives - Urban Geo Analytics</title>
	<atom:link href="https://urbangeoanalytics.com/tag/open-source/feed/" rel="self" type="application/rss+xml" />
	<link>https://urbangeoanalytics.com/tag/open-source/</link>
	<description>Spatial Analysis, GeoAI &#38; Machine Learning</description>
	<lastBuildDate>Mon, 17 Aug 2026 18:44:00 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1</generator>

<image>
	<url>https://urbangeoanalytics.com/wp-content/uploads/2025/11/cropped-logo-urban-geo_512-32x32.png</url>
	<title>Open Source Archives - Urban Geo Analytics</title>
	<link>https://urbangeoanalytics.com/tag/open-source/</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>UVLM v4.0.0 — Gemma 4, the Transformers 5 Migration, and Why This One Is a Major Version</title>
		<link>https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/</link>
					<comments>https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Mon, 17 Aug 2026 07:45:10 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Package]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[gemma]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=3029</guid>

					<description><![CDATA[<p>Highlights  New model family: Gemma 4 (Google DeepMind, released April 2026) joins as the fifth family — E2B, E4B, and 12B Instruct, bringing the registry to 24 checkpoints Breaking change, done honestly: UVLM now requires Transformers ≥ 5.15; all four existing families were re-validated on GPU before release, and v3.2.0 remains installable  [...]</p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/">UVLM v4.0.0 — Gemma 4, the Transformers 5 Migration, and Why This One Is a Major Version</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<p><div class="fusion-fullwidth fullwidth-box fusion-builder-row-1 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-0 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-1 hover-type-none"><img fetchpriority="high" decoding="async" width="1536" height="1024" title="UVLM 4.0.0" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0.png" alt class="img-responsive wp-image-3043" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/UVLM-4.0.0.png 1536w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-1 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-1 fusion-text-no-margin" style="--awb-margin-bottom:-10px;"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-2" style="--awb-margin-top:-20px;"><ul>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>New model family:</strong> Gemma 4 (Google DeepMind, released April 2026) joins as the fifth family — E2B, E4B, and 12B Instruct, bringing the registry to 24 checkpoints</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Breaking change, done honestly:</strong> UVLM now requires Transformers ≥ 5.15; all four existing families were re-validated on GPU before release, and v3.2.0 remains installable for Transformers 4.x environments</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Measured, not guessed:</strong> Gemma 4&#8217;s &#8220;effective parameters&#8221; hide ~10–16 GB raw checkpoints — this release documents exactly what runs on an 8 GB GPU, and how</li>
</ul>
</div><div class="fusion-title title fusion-title-1 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Why a major version?</span></h2></div><div class="fusion-text fusion-text-3 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr"><a class="keychainify-checked" href="https://github.com/perezjoan/UVLM">UVLM</a> has followed one rule since the package release: a new model family is a minor version, because it breaks nothing. v4.0.0 breaks that streak for a reason we could not code around: <strong>Gemma 4 does not exist in any Transformers 4.x release.</strong> We verified this empirically — 4.57.6 is the final version of the 4.x line, and it does not register the <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">gemma4</code> architecture; support begins in the 5.x line. Adopting the family therefore means lifting the <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">transformers &lt; 5.0.0</code> cap that UVLM has carried since v3.0.1.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">That cap was not decoration. It existed because early Transformers 5.x releases crashed Qwen2.5-VL at load time with a weight-conversion error. So before this release, all four existing families — LLaVA-NeXT, Qwen2.5-VL, Qwen3-VL, InternVL3.5 — were re-validated on GPU under Transformers 5.15, in 4-bit, on real inference tasks. The historical Qwen2.5-VL crash is <strong>confirmed fixed</strong>: the 7B model loads and answers correctly at full speed. That validation is what makes this a release rather than a gamble.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">If your environment must stay on Transformers 4.x, nothing is taken from you: <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">pip install git+https://github.com/perezjoan/UVLM.git@v3.2.0</code> pins the last 4.x-compatible release, permanently.</p>
</div><div class="fusion-title title fusion-title-2 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What is Gemma 4?</span></h2></div><div class="fusion-text fusion-text-4 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Gemma 4 is Google DeepMind&#8217;s latest open multimodal generation, released in April 2026 under Apache 2.0. UVLM v4.0.0 integrates three Instruct checkpoints:</p>
</div>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Model</th>
<th align="left"> Parameters</th>
<th align="left"> Raw checkpoint</th>
<th align="left"> Runs on</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Gemma 4 E2B Instruct</td>
<td align="left"> ~2B effective</td>
<td align="left"> ~10 GB</td>
<td align="left"> 8 GB GPU in FP16 with CPU offload</td>
</tr>
<tr>
<td align="left">Gemma 4 E4B Instruct</td>
<td align="left"> ~4B effective</td>
<td align="left"> ~16 GB</td>
<td align="left">Larger-VRAM environments (Colab A100/L4)</td>
</tr>
<tr>
<td align="left">Gemma 4 12B Instruct</td>
<td align="left"> 12B</td>
<td align="left"> ~24 GB</td>
<td align="left">Larger-VRAM environments (Colab A100/L4)</td>
</tr>
</tbody>
</table>
</div>
<div class="fusion-text fusion-text-5 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Notice the third column, because it is this release&#8217;s most useful finding. E2B and E4B are <strong>&#8220;effective&#8221;-parameter models</strong>: Per-Layer Embeddings give them the <em>compute</em> profile of a 2B/4B model, but the embedding tables push the <em>raw</em> checkpoint far beyond what the name suggests. A &#8220;2B&#8221; model that downloads 10 GB of weights behaves very differently from Qwen3-VL 2B&#8217;s genuinely small footprint — and honest benchmarking infrastructure should say so, with numbers.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Technically, Gemma 4 shares the tokenizing-chat-template pipeline introduced with InternVL3.5, with one addition: Gemma 4 models other than E2B/E4B wrap their output in thought-channel tags even when thinking is disabled, and the backend strips them automatically. As always, the family appeared in the notebook selector with zero interface changes.</p>
</div><div class="fusion-title title fusion-title-3 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What actually runs on a laptop GPU</span></h2></div><div class="fusion-text fusion-text-6 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>We validated Gemma 4 on an 8 GB RTX 5060, and the result inverts the usual intuition: <strong>FP16 is the low-memory mode.</strong> In FP16, the compute-heavy layers stay on the GPU while the PLE embedding tables — lookup-only structures designed to live off-accelerator — offload to system RAM in half precision. Measured: about 17 seconds per image. Slow, but fully functional.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-2 hover-type-none"><img decoding="async" width="1215" height="641" title="gemma illustration" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration.png" alt class="img-responsive wp-image-3039" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-200x106.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-400x211.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-600x317.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-800x422.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration-1200x633.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/gemma-illustration.png 1215w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div><div class="fusion-text fusion-text-7 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>4-bit quantization — normally the memory-saver — fails here, for a subtle reason: offloaded modules are kept in FP32, roughly doubling the RAM requirement, and any further spill to disk is unsupported by bitsandbytes. Rather than leave users with a 200-line traceback, v4.0.0 detects this case and raises a two-sentence error recommending FP16. To support all of this, the loader gained general CPU-offload capability for oversized checkpoints — a change that only <em>permits</em> offload: models that fit entirely on the GPU are placed exactly as before.</p>
</div><div class="fusion-title title fusion-title-4 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Also in this release</span></h2></div><div class="fusion-text fusion-text-8 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>A failed model load in the notebooks now invalidates the previously loaded model, so a batch run after a failed load errors out loudly instead of silently benchmarking the wrong checkpoint — a trap we fell into ourselves during validation, and one that the per-model output filenames from v3.2.0 caught. InternVL3.5 users on Transformers 5 will see a harmless &#8220;tied weights&#8221; warning caused by an upstream config inconsistency; Transformers resolves it correctly.</p>
</div><div class="fusion-title title fusion-title-5 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Getting started</span></h2></div><div class="fusion-text fusion-text-9" style="--awb-margin-top:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="Bash" data-enlighter-title="Bash">pip install --upgrade --force-reinstall git+https://github.com/perezjoan/UVLM.git</pre>
</div><div class="fusion-text fusion-text-10 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-bottom:25px;"><p>Note: no <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">--no-deps</code> this time — the whole point of the upgrade is that pip pulls Transformers 5.15 for you. Colab users get v4.0.0 automatically on their next session, since the notebook always installs the latest version; the re-validation above is what makes that automatic jump safe. The three-block workflow, consensus validation, chain-of-thought mode, and truncation detection all work with Gemma 4 out of the box.</p>
</div><div class="fusion-title title fusion-title-6 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Where UVLM stands</span></h2></div><div class="fusion-text fusion-text-11 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Five months ago, UVLM was a two-family package. It now supports <strong>five families and 24 checkpoints from 1B to 110B parameters</strong> — LLaVA-NeXT, Qwen2.5-VL, Qwen3-VL, InternVL3.5, Gemma 4 — behind one interface, one prompt format, one evaluation protocol. Three families were added in three releases without a single notebook edit, each validated on hardware before shipping. That is the registry we will be benchmarking against in upcoming applied work — more on that soon.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Full change log in <a class="keychainify-checked" href="https://github.com/perezjoan/UVLM/blob/main/VERSIONS.txt">VERSIONS.txt</a> · Source and releases on GitHub · If you use UVLM in research, please cite our Software paper.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-2 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-12"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--1" data-awb-toc-id="1" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-3 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div><div class="fusion-fullwidth fullwidth-box fusion-builder-row-2 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"></div></div></p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/">UVLM v4.0.0 — Gemma 4, the Transformers 5 Migration, and Why This One Is a Major Version</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/uvlm-4-0-0-gemma-4-transformers-5/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>UVLM v3.2.0 — InternVL3.5 Joins the Registry, With Zero Notebook Changes</title>
		<link>https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/</link>
					<comments>https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Wed, 12 Aug 2026 08:14:01 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Package]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[InternVL]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2982</guid>

					<description><![CDATA[<p>UVLM v3.2.0 adds InternVL3.5 (1B–38B, six checkpoints): 21 open VLM checkpoints across 4 families, one Python interface. The new family appeared in the notebooks without a single notebook edit — plus per-model output files for cleaner benchmarking.</p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/">UVLM v3.2.0 — InternVL3.5 Joins the Registry, With Zero Notebook Changes</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-3 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-3 fusion_builder_column_1_1 1_1 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:100%;--awb-margin-top-large:0px;--awb-spacing-right-large:1.92%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:1.92%;--awb-width-medium:100%;--awb-order-medium:0;--awb-spacing-right-medium:1.92%;--awb-spacing-left-medium:1.92%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-4 hover-type-none"><img decoding="async" width="1536" height="1024" title="uvlm3.2.0 illustration" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration.png" alt class="img-responsive wp-image-2991" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm3.2.0-illustration.png 1536w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-4 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-title title fusion-title-7 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Highlights</span></h2></div><div class="fusion-text fusion-text-13"><ul>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>New model family:</strong> InternVL3.5 (OpenGVLab, released August 2025) joins LLaVA-NeXT, Qwen2.5-VL, and Qwen3-VL — six checkpoints from 1B to 38B</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Zero notebook changes:</strong> the new family appeared in the selector automatically — the extensibility promise from v3.1.0, kept</li>
<li class="font-claude-response-body whitespace-normal break-words pl-2"><strong>Per-model output files:</strong> each checkpoint now writes its own CSV, so resume mode can never mix results from different models</li>
</ul>
</div><div class="fusion-title title fusion-title-8 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">1. What is InternVL3.5?</span></h2></div><div class="fusion-text fusion-text-14 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">In <a class="keychainify-checked" href="https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/">v3.1.0</a> we added Qwen3-VL and made a promise: thanks to the new <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">FAMILY_GROUPS</code> registry, future model families would appear in the notebooks automatically, with no interface edits at all. Version 3.2.0 is that promise kept. <strong>InternVL3.5</strong> — the latest generation of OpenGVLab&#8217;s InternVL line, released in August 2025 — is now the fourth family in the registry, and neither notebook changed by a single line to display it.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">UVLM integrates the six Transformers-native <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">-HF</code> checkpoints, which run through the standard Transformers stack without any custom remote code. None of them is gated: no Hugging Face token required.</p>
</div>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Parameters</th>
<th align="left">VRAM (4-bit)</th>
<th align="left"> Typical hardware</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">InternVL3.5 1B</td>
<td align="left"> 1B</td>
<td align="left">~1 GB</td>
<td align="left">Any modern laptop GPU, free Colab T4</td>
</tr>
<tr>
<td align="left">InternVL3.5 2B</td>
<td align="left">2B</td>
<td align="left">~2 GB</td>
<td align="left">Any modern laptop GPU, free Colab T4</td>
</tr>
<tr>
<td align="left">InternVL3.5 4B</td>
<td align="left">4B</td>
<td align="left">~3 GB</td>
<td align="left">T4, RTX 3060</td>
</tr>
<tr>
<td align="left">InternVL3.5 8B</td>
<td align="left">8B</td>
<td align="left">~6 GB</td>
<td align="left">T4, RTX 4060/5060</td>
</tr>
<tr>
<td align="left">InternVL3.5 14B</td>
<td align="left">14B</td>
<td align="left">~9 GB</td>
<td align="left">L4, RTX 4070</td>
</tr>
<tr>
<td align="left">InternVL3.5 38B</td>
<td align="left">38B</td>
<td align="left">~22 GB</td>
<td align="left">A100, RTX 4090</td>
</tr>
</tbody>
</table>
</div>
<div class="fusion-text fusion-text-15 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The registry now totals <strong>21 checkpoints across 4 families</strong>, from 1B to 110B parameters — and the 1B entry replaces Qwen3-VL 2B as the smallest model UVLM has ever supported.</p>
</div><div class="fusion-title title fusion-title-9 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">2. A genuinely different pipeline</span></h2></div><div class="fusion-text fusion-text-16 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">InternVL3.5 is not a variation on the Qwen conventions — it uses the standard Transformers pattern in which the <strong>chat template tokenizes directly</strong> (<code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">apply_chat_template(tokenize=True)</code>), the generated tokens are sliced off after the prompt, and only the generated portion is decoded. That makes it the third distinct inference path in UVLM, alongside LLaVA&#8217;s string-based cleaning and Qwen&#8217;s separate vision preprocessing with token trimming. As always, all three converge at the same unified response parser — from the user&#8217;s side, InternVL3.5 is simply one more family in the dropdown.</p>
</div><div class="fusion-image-element " style="--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-5 hover-type-none"><img decoding="async" width="1308" height="644" title="internvl uvlm" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm.png" alt class="img-responsive wp-image-2983" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-200x98.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-400x197.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-600x295.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-800x394.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm-1200x591.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/internvl-uvlm.png 1308w" sizes="(max-width: 640px) 100vw, 1200px" /></span></div><div class="fusion-text fusion-text-17 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Loading follows the same BF16-aware logic introduced in v3.1.0: BF16 automatically on GPUs with native support (RTX 30-series and newer, L4, A100), FP16 fallback otherwise.</p>
</div><div class="fusion-title title fusion-title-10 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">3. One honest bug fix: per-model output files</span></h2></div><div class="fusion-text fusion-text-18 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">While validating the new backend, we caught a fossil from UVLM&#8217;s two-backend era: the notebooks used a hardcoded rule that sent every non-Qwen2.5 model&#8217;s results to <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">Score_Analysis_LLaVA.csv</code>. With four families, that meant different models could silently append into the same CSV — and resume mode could not tell them apart. As of v3.2.0, <strong>output filenames are derived from the loaded checkpoint</strong> (e.g. <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">Score_Analysis_InternVL3_5-8B-HF.csv</code>), so each model writes its own file and resume mode and schema upgrades are per-model by construction. If you benchmark several models on the same image folder, this is the release that keeps your results honest.</p>
</div><div class="fusion-title title fusion-title-11 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">4. Getting started</span></h2></div><div class="fusion-text fusion-text-19 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:5px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Nothing changes in the workflow — install (or upgrade) and the new family is there:</p>
</div><div class="fusion-text fusion-text-20 fusion-text-no-margin" style="--awb-margin-top:5px;--awb-margin-bottom:5px;"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="Bash1" data-enlighter-title="Bash">pip install --upgrade --force-reinstall --no-deps
git+https://github.com/perezjoan/UVLM.git</pre>
</div><div class="fusion-text fusion-text-21 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:5px;--awb-margin-bottom:5px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Or open the Colab notebook — it always installs the latest version automatically. The three-block workflow (load → configure tasks → run batch), consensus validation, chain-of-thought mode, and truncation detection all work with InternVL3.5 out of the box. No dependency changes since v3.1.0. Tested locally on Windows 11 with an RTX 5060 laptop GPU, where the 1B model loads in about 15 seconds once cached.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">One field note from validation, and a nice illustration of why UVLM separates format reliability from accuracy: under temperature sampling, InternVL3.5 1B answered a counting task with &#8220;There are two vehicles in the picture&#8221; — correct, but unparseable as an integer, so it was recorded as NA by design. Under greedy decoding with a strict format instruction, the same model returned a clean integer. Small models follow instructions best when you ask firmly and decode greedily.</p>
</div><div class="fusion-title title fusion-title-12 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">5. What&#8217;s next</span></h2></div><div class="fusion-text fusion-text-22 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr">The third family addition promised in v3.1.0 — the <strong>Gemma</strong> multimodal line — is coming next, and it will be a bigger step than a minor version: Gemma 4 requires the Transformers v5 line, which means UVLM&#8217;s next release will be a <strong>major version</strong> with a documented migration. Same discipline as always: one backend at a time, validated before released.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr">Full change log in VERSIONS.txt · Source and releases on <a class="keychainify-checked" href="https://github.com/perezjoan/UVLM">GitHub</a> · If you use UVLM in research, please cite our <a class="keychainify-checked" href="https://www.mdpi.com/2674-113X/5/3/30">Software paper</a>.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-5 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-23"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--2" data-awb-toc-id="2" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-6 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/">UVLM v3.2.0 — InternVL3.5 Joins the Registry, With Zero Notebook Changes</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/uvlm-3-2-0-internvl-backend/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>UVLM v3.1.0 — Qwen3-VL Joins the Registry, With Family-Based Model Selection</title>
		<link>https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/</link>
					<comments>https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Mon, 10 Aug 2026 07:46:22 +0000</pubDate>
				<category><![CDATA[Advanced]]></category>
		<category><![CDATA[Package]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[AI]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[Qwen]]></category>
		<category><![CDATA[UVLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2898</guid>

					<description><![CDATA[<p>UVLM v3.1.0 adds a third model family, Qwen3-VL (2B–32B Instruct), bringing the registry to 15 checkpoints. The notebooks gain a two-level family/model selector, the loader picks BF16 automatically on capable GPUs, and the smallest new model runs in about 2 GB of VRAM. Same three-block workflow, same prompts, one more family to compare.</p>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/">UVLM v3.1.0 — Qwen3-VL Joins the Registry, With Family-Based Model Selection</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-4 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-6 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-24"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-25" style="--awb-margin-top:-30px;"><ul>
<li><strong data-start="64" data-end="88">New model family:</strong>Qwen3-VL (released from September 2025) joins LLaVA-NeXT and Qwen2.5-VL</li>
<li><strong>Two-level model selection</strong>: pick the family first, the model list refreshes automatically</li>
<li><strong>Lightest model yet</strong>: Qwen3-VL 2B runs in ~2 GB of VRAM with 4-bit quantization</li>
</ul>
</div><div class="fusion-title title fusion-title-13 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What is Qwen3-VL?</span></h2></div><div class="fusion-text fusion-text-26 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>UVLM was built around one idea: compare Vision-Language Models across architectures using <strong>identical prompts and evaluation protocols</strong>, without writing model-specific code. Until now that meant two families: LLaVA-NeXT and Qwen2.5-VL. Version 3.1.0 adds a third: <strong>Qwen3-VL</strong>, the successor to the Qwen2.5-VL family that anchored our published benchmark.</p>
<p>Qwen3-VL is the latest vision-language generation from Alibaba&#8217;s Qwen team, first released in <strong>September 2025</strong> with the 235B-A22B flagship, followed shortly after by the compact dense checkpoints that matter for most research budgets. UVLM v3.1.0 integrates the four dense Instruct sizes:</p>
</div>
<div class="table-1">
<table width="100%">
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Parameters</th>
<th align="left"> VRAM (4-bit)</th>
<th align="left">Typical hardware</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Qwen3-VL 2B Instruct</td>
<td align="left">2B</td>
<td align="left"> ~2 GB</td>
<td align="left">Any modern laptop GPU, free Colab T4</td>
</tr>
<tr>
<td align="left">Qwen3-VL 4B Instruct</td>
<td align="left">4B</td>
<td align="left"> ~3 GB</td>
<td align="left"> T4, RTX 3060</td>
</tr>
<tr>
<td align="left">Qwen3-VL 8B Instruct</td>
<td align="left">8B</td>
<td align="left"> ~6 GB</td>
<td align="left"> T4, RTX 4060/5060</td>
</tr>
<tr>
<td align="left">Qwen3-VL 32B Instruct</td>
<td align="left">32B</td>
<td align="left"> ~20 GB</td>
<td align="left"> A100, RTX 4090</td>
</tr>
</tbody>
</table>
</div>
<div class="fusion-text fusion-text-27 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">The registry now totals <strong>15 checkpoints across 3 families</strong>, from 2B to 110B parameters — and the 2B entry is the smallest model UVLM has ever supported, which makes it an interesting new baseline for large-scale, low-cost batch analysis.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="34:1-34:532;2755-3286">Technically, Qwen3-VL keeps the Qwen inference conventions (chat template → separate vision preprocessing → generation → token trimming), so it plugs into UVLM&#8217;s existing Qwen pipeline. What changes under the hood: the model loads through Transformers&#8217; generic <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">AutoModelForImageTextToText</code> class, requires <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">transformers ≥ 4.57</code> and <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">qwen-vl-utils ≥ 0.0.14</code>, and resizes images to multiples of 32 pixels rather than 28. All of this is handled inside the package — from the user&#8217;s side, it is simply one more family in the dropdown.</p>
</div><div class="fusion-title title fusion-title-14 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Pick the family, then the model</span></h2></div><div class="fusion-text fusion-text-28 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">With three families and fifteen checkpoints, a single flat dropdown was getting crowded. Both notebooks (Colab and local) now use a <strong>two-level selector</strong>: choose the family first — LLaVA-NeXT, Qwen2.5-VL, or Qwen3-VL — and the model list refreshes automatically.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-7" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-7 hover-type-none"><img decoding="async" width="1140" height="454" title="uvlm family" src="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family.png" alt class="img-responsive wp-image-2900" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-200x80.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-400x159.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-600x239.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family-800x319.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/08/uvlm-family.png 1140w" sizes="(max-width: 640px) 100vw, 1140px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">UVLM v3.1.0: Two-level model selection</div></div></div></div><div class="fusion-text fusion-text-29 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">The selector is built from a new <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">FAMILY_GROUPS</code> mapping in the registry, which means future families will appear in the widgets automatically, with no notebook edits at all.</p>
</div><div class="fusion-title title fusion-title-15 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Smarter precision handling</span></h2></div><div class="fusion-text fusion-text-30 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">Qwen3-VL checkpoints are trained in BF16. On GPUs with native BF16 support (RTX 30-series and newer, L4, A100), the loader now selects <strong>BF16 automatically</strong>, falling back to FP16 on older cards and FP32 on CPU. If you followed our earlier benchmark work, you may remember the FP16 numerical-overflow crashes we documented with BF16-trained checkpoints on T4 hardware — this release is the first step toward closing that class of problem at the loader level.</p>
</div><div class="fusion-title title fusion-title-16 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">Quality-of-life fixes</span></h2></div><div class="fusion-text fusion-text-31 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">Local Jupyter users get a long-overdue improvement: the model-loading progress (download bars, device map, timings) is now displayed in a <strong>log area under the Load button</strong>. Previously, output emitted inside the widget callback was silently swallowed in local Jupyter — Colab was never affected. The release also silences the <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">torch_dtype</code> deprecation warnings from recent Transformers versions and synchronizes the package version metadata.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="54:1-54:84;4795-4878">Nothing changes in the workflow — install (or upgrade) and the new family is there:</p>
</div><div class="fusion-text fusion-text-32 fusion-text-no-margin" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><pre class="EnlighterJSRAW" data-enlighter-language="bash" data-enlighter-theme="dracula" data-enlighter-group="bash1" data-enlighter-title="bash">pip install --upgrade --force-reinstall --no-deps git+https://github.com/perezjoan/UVLM.git</pre>
</div><div class="fusion-text fusion-text-33 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="32:1-32:240;2514-2753">Or open the <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://colab.research.google.com/github/perezjoan/UVLM/blob/main/notebooks/UVLM_colab.ipynb">Colab notebook</a> — it always installs the latest version automatically. The three-block workflow (load → configure tasks → run batch), consensus validation, chain-of-thought mode, and truncation detection all work with Qwen3-VL out of the box. Tested locally on Windows 11 with an RTX 5060 laptop GPU, where the 2B model loads in well under a minute once cached.</p>
</div><div class="fusion-title title fusion-title-17 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:0px;--awb-font-size:35px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;font-size:1em;--fontSize:35;line-height:var(--awb-typography1-line-height);"><span style="font-weight: 400;">What’s next</span></h2></div><div class="fusion-text fusion-text-34 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="64:1-64:266;5471-5736">v3.1.0 is the first of a series of family additions. Next on the roadmap: <strong>InternVL3.5</strong> (via the Transformers-native <code class="bg-text-200/5 border border-0.5 border-border-300 text-danger-000 whitespace-pre-wrap rounded-&#091;0.4rem&#093; px-1 py-px text-&#091;0.9rem&#093;">-HF</code> checkpoints) and the <strong>Gemma</strong> multimodal line. Each family will land as its own validated release — same discipline, one backend at a time.</p>
<p class="font-claude-response-body break-words whitespace-normal" dir="ltr" data-sourcepos="66:1-66:265;5738-6002">Full change log in <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://github.com/perezjoan/UVLM/blob/main/VERSIONS.txt">VERSIONS.txt</a> · Source and releases on <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://github.com/perezjoan/UVLM">GitHub</a> · If you use UVLM in research, please cite our <a class="underline underline-offset-2 decoration-1 decoration-current/40 hover:decoration-current focus:decoration-current keychainify-checked" href="https://www.mdpi.com/2674-113X/5/3/30">Software paper</a>.</p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-7 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-35"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--3" data-awb-toc-id="3" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-8 hover-type-zoomout"><img decoding="async" width="1536" height="1024" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png" alt class="img-responsive wp-image-1688" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/blog-lvl3.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/">UVLM v3.1.0 — Qwen3-VL Joins the Registry, With Family-Based Model Selection</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/uvlm-3-1-0-qwen3-vl-backend/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>Introducing UVLM: A Free Tool to Compare AI Models That Understand Images</title>
		<link>https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/</link>
					<comments>https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/#respond</comments>
		
		<dc:creator><![CDATA[Joan Perez]]></dc:creator>
		<pubDate>Tue, 17 Mar 2026 14:23:58 +0000</pubDate>
				<category><![CDATA[Intermediate]]></category>
		<category><![CDATA[Python]]></category>
		<category><![CDATA[Vision Language Model]]></category>
		<category><![CDATA[Benchmarking]]></category>
		<category><![CDATA[Chain-of-Thought]]></category>
		<category><![CDATA[Google Colab]]></category>
		<category><![CDATA[Image Analysis]]></category>
		<category><![CDATA[Llava]]></category>
		<category><![CDATA[Multimodal AI]]></category>
		<category><![CDATA[Open Source]]></category>
		<category><![CDATA[Qwen]]></category>
		<category><![CDATA[UVLM]]></category>
		<category><![CDATA[VLM]]></category>
		<guid isPermaLink="false">https://urbangeoanalytics.com/?p=2356</guid>

					<description><![CDATA[<p>UVLM is a free, open-source tool for loading, testing, and comparing Vision-Language Models on custom image analysis tasks. Running entirely in Google Colab, it lets researchers and practitioners benchmark multiple AI models using the same prompts and images — no coding, no GPU ownership, no model-specific pipelines. This post explains what VLMs are, why comparing them matters, and how to get started in five minutes.</p>
<p>The post <a href="https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/">Introducing UVLM: A Free Tool to Compare AI Models That Understand Images</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></description>
										<content:encoded><![CDATA[<div class="fusion-fullwidth fullwidth-box fusion-builder-row-5 fusion-flex-container has-pattern-background has-mask-background nonhundred-percent-fullwidth non-hundred-percent-height-scrolling" style="--awb-border-radius-top-left:0px;--awb-border-radius-top-right:0px;--awb-border-radius-bottom-right:0px;--awb-border-radius-bottom-left:0px;--awb-flex-wrap:wrap;" id="contenu" ><div class="fusion-builder-row fusion-row fusion-flex-align-items-flex-start fusion-flex-content-wrap" style="max-width:1248px;margin-left: calc(-4% / 2 );margin-right: calc(-4% / 2 );"><div class="fusion-layout-column fusion_builder_column fusion-builder-column-8 fusion_builder_column_3_4 3_4 fusion-flex-column" style="--awb-bg-size:cover;--awb-width-large:75%;--awb-margin-top-large:0px;--awb-spacing-right-large:2.56%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:2.56%;--awb-width-medium:75%;--awb-order-medium:0;--awb-spacing-right-medium:2.56%;--awb-spacing-left-medium:2.56%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;" id="contenu" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-9" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-9 hover-type-none"><img decoding="async" width="1536" height="595" title="uvlm" src="https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm.png" alt class="img-responsive wp-image-2342" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-200x77.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-400x155.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-600x232.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-800x310.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm-1200x465.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/uvlm.png 1536w" sizes="(max-width: 640px) 100vw, 1200px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">uvlm</div></div></div></div><div class="fusion-text fusion-text-36"><h5><strong>Highlights</strong></h5>
</div><div class="fusion-text fusion-text-37" style="--awb-margin-top:-30px;"><ul>
<li><strong>New open-source release: UVLM v2.2.2</strong> — compare Vision-Language Models from a single notebook</li>
<li><strong>11 AI models</strong>, 5 analysis tasks, 120 test images — all benchmarked with one tool</li>
<li><strong>No coding, no installation</strong> — runs in Google Colab with a free account</li>
</ul>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-text fusion-text-38 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Imagine you have thousands of street photographs and you need to answer the same questions about each one: how many cars are parked? Is there a sidewalk? How long is the building frontage? Hiring someone to go through every image manually would take weeks. Training a custom computer vision model would take months. But what if you could simply ask an AI model these questions in plain English — and get structured, usable answers back?</p>
<p>That is exactly what Vision-Language Models do. And today, we are releasing UVLM — an open-source tool that makes it easy to load, test, and compare these models, all from a single notebook in your browser.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-18 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">What Are Vision-Language Models?</h2></div><div class="fusion-text fusion-text-39 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Vision-Language Models (VLMs) are AI systems that can look at an image and answer questions about it in natural language. Unlike traditional computer vision, which requires training a separate model for every task (one for counting cars, another for detecting sidewalks, a third for classifying buildings), a VLM handles all of these through text prompts. You write a question, attach a photo, and the model responds.</p>
<p>For example, you can ask a VLM: “Count all motor vehicles visible in this image” and it will answer “3”. You can ask the same model “Is there a sidewalk along the street frontage?” and it will answer “yes”. You can even ask it to estimate the length of a building facade in meters — a task that requires the model to identify reference objects (like parked cars), estimate their size, and reason about perspective. All of this from a single model, with no retraining and no labelled dataset.</p>
<p>The catch is that there are many VLM families available (LLaVA, Qwen, InternVL, BLIP-2, and more), and each one works differently under the hood. They use different image encoders, different tokenisation strategies, and different code to run. If you want to know which model is best for your specific task, you normally have to write separate code for each one — a tedious and error-prone process.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-19 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">This Is the Problem UVLM Solves</h2></div><div class="fusion-text fusion-text-40 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>UVLM (Universal Vision-Language Model Loader) is a free, open-source tool that lets you load, configure, and compare multiple VLM architectures using the same prompts and the same evaluation protocol — without writing any model-specific code. It runs entirely in Google Colab, which means you do not need to install anything on your computer or own a GPU. A free Google account is all you need.</p>
<p>The idea is simple: you pick a model from a dropdown menu, type your analysis questions into a form, point the tool at a folder of images, and hit run. UVLM handles all the technical details — the processor classes, the tokenisation, the generation settings, the output parsing — and delivers a clean CSV file with one row per image and one column per task. If you want to try a different model, you just switch the dropdown and run again. Same prompts, same images, same output format. Now you can compare.</p>
</div><div class="fusion-image-element awb-imageframe-style awb-imageframe-style-below awb-imageframe-style-10" style="text-align:center;--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--body_typography-font-family);--awb-caption-title-font-weight:var(--body_typography-font-weight);--awb-caption-title-font-style:var(--body_typography-font-style);--awb-caption-title-size:var(--body_typography-font-size);--awb-caption-title-transform:var(--body_typography-text-transform);--awb-caption-title-line-height:var(--body_typography-line-height);--awb-caption-title-letter-spacing:var(--body_typography-letter-spacing);"><span class=" fusion-imageframe imageframe-none imageframe-10 hover-type-none"><img decoding="async" width="1190" height="823" title="image1" src="https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1.png" alt class="img-responsive wp-image-2319" srcset="https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-200x138.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-400x277.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-600x415.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1-800x553.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2026/03/image1.png 1190w" sizes="(max-width: 640px) 100vw, 1190px" /></span><div class="awb-imageframe-caption-container" style="text-align:center;"><div class="awb-imageframe-caption"><div class="awb-imageframe-caption-title">The 3 blocks structure of UVLM Loader</div></div></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-20 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">A Practical Example: Scoring 120 Street Photographs</h2></div><div class="fusion-text fusion-text-41 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>To demonstrate what UVLM can do, we benchmarked 8 different models on 120 street-level photographs of French urban frontages. Each image was analysed on five tasks: counting vehicles, detecting sidewalks, counting pedestrian entrances, estimating the street frontage length in meters, and classifying the vegetation type. That is 16 model configurations (each model tested in standard and advanced reasoning modes), 120 images, and 5 tasks per image — all processed and compared through UVLM.</p>
<p>The results were revealing. The largest model (LLaVA 34B, with 34 billion parameters) actually ranked last overall. A much smaller model (LLaVA Vicuna 7B) outperformed it significantly and ran on a free Google Colab GPU. The best overall results came from Qwen 32B with chain-of-thought reasoning enabled, which achieved 88% proximity to human expert annotations across all five tasks. Without UVLM, discovering these differences would have required writing and debugging eight separate inference pipelines.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-21 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">Who Is UVLM For?</h2></div><div class="fusion-text fusion-text-42 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>UVLM was designed for anyone who works with images and wants to extract structured information from them at scale — without becoming a machine learning engineer. If you are an urban planner evaluating streetscape quality across a city, UVLM lets you score thousands of street photographs using natural language prompts. If you are an environmental researcher classifying vegetation from field photographs, UVLM lets you test which AI model gives the most reliable results for your specific classification scheme. If you are an infrastructure inspector processing damage assessment photographs, UVLM lets you set up automated counting and scoring tasks and run them across your entire image archive.</p>
<p>The tool is also valuable for AI researchers who need a controlled benchmarking environment. Because UVLM ensures that every model receives exactly the same prompt and is evaluated with the same metrics, it produces fair, reproducible comparisons. The consensus validation feature (running each task multiple times and taking a majority vote) addresses the inherent randomness of AI outputs, and the truncation detection feature flags when a model’s response was cut off before it could finish — a common but often invisible source of errors.</p>
</div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:25px;margin-bottom:25px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-title title fusion-title-22 fusion-sep-none fusion-title-text fusion-title-size-two" style="--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2 class="fusion-title-heading title-heading-left fusion-responsive-typography-calculated" style="margin:0;--fontSize:48;line-height:var(--awb-typography1-line-height);">How to Get Started</h2></div><div class="fusion-text fusion-text-43 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><p>Getting started takes about five minutes. Open the UVLM notebook from GitHub (the link is below), connect to a GPU runtime in Google Colab, and run the first block to load a model. The second block gives you a form where you type your analysis questions — no coding required. The third block processes your images and saves the results as a CSV file on your Google Drive.</p>
<p>The tool currently supports 11 model checkpoints from two major families (LLaVA-NeXT and Qwen2.5-VL), ranging from 3 billion to 110 billion parameters. Models up to 34B can run on a single free-tier Colab GPU with 4-bit quantisation. Advanced features include consensus validation (2–5 runs per task with majority voting), chain-of-thought reasoning for complex tasks, and automatic truncation detection.</p>
<p>UVLM is released under the Apache 2.0 open-source licence. You can use it, modify it, and build on it for any purpose — academic or commercial.</p>
</div><div class="fusion-text fusion-text-44 fusion-text-no-margin" style="--awb-content-alignment:justify;--awb-margin-top:25px;--awb-margin-bottom:25px;"><h2>Links</h2>
<p><strong>Source code: </strong><a class="keychainify-checked" href="https://github.com/perezjoan/UVLM">github.com/perezjoan/UVLM</a></p>
<p><strong>Paper: </strong><a class="keychainify-checked" href="https://arxiv.org/abs/2603.13893">arXiv preprint — Perez &amp; Fusco (2026)</a></p>
<p><strong>UVLM page on this site: </strong><a class="keychainify-checked" href="https://urbangeoanalytics.com/algorithms-softwares/uvlm-universal-vision-language-model-loader/">urbangeoanalytics.com › Softwares &amp; Algorithms › UVLM</a></p>
<p><strong>Benchmark dataset: </strong><a class="keychainify-checked" href="https://zenodo.org/records/18959690">Zenodo — 120 street-view images</a></p>
<h2>Citation</h2>
<p>If you use UVLM in your work, please cite:</p>
<p><em>Perez, J. &amp; Fusco, G. (2026). UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking. arXiv:2603.13893</em></p>
</div></div></div><div class="fusion-layout-column fusion_builder_column fusion-builder-column-9 awb-sticky awb-sticky-medium awb-sticky-large fusion_builder_column_1_4 1_4 fusion-flex-column" style="--awb-padding-top:20px;--awb-padding-right:20px;--awb-padding-bottom:20px;--awb-padding-left:20px;--awb-bg-size:cover;--awb-border-color:var(--awb-color6);--awb-border-style:solid;--awb-width-large:25%;--awb-margin-top-large:0px;--awb-spacing-right-large:7.68%;--awb-margin-bottom-large:20px;--awb-spacing-left-large:7.68%;--awb-width-medium:25%;--awb-order-medium:0;--awb-spacing-right-medium:7.68%;--awb-spacing-left-medium:7.68%;--awb-width-small:100%;--awb-order-small:0;--awb-spacing-right-small:1.92%;--awb-spacing-left-small:1.92%;--awb-sticky-offset:150px;" data-scroll-devices="small-visibility,medium-visibility,large-visibility"><div class="fusion-column-wrapper fusion-column-has-shadow fusion-flex-justify-content-flex-start fusion-content-layout-column"><div class="fusion-text fusion-text-45"><p><span style="color: #143c4e;"><strong>Table of contents</strong></span></p>
</div><div class="awb-toc-el awb-toc-el--4" data-awb-toc-id="4" data-awb-toc-options="{&quot;allowed_heading_tags&quot;:{&quot;h2&quot;:0},&quot;ignore_headings&quot;:&quot;&quot;,&quot;ignore_headings_words&quot;:&quot;&quot;,&quot;enable_cache&quot;:&quot;no&quot;,&quot;highlight_current_heading&quot;:&quot;yes&quot;,&quot;hide_hidden_titles&quot;:&quot;no&quot;,&quot;limit_container&quot;:&quot;page_content&quot;,&quot;select_custom_headings&quot;:&quot;.contenu H2, .contenu H3&quot;,&quot;icon&quot;:&quot;fa-flag fas&quot;,&quot;counter_type&quot;:&quot;none&quot;}" style="--awb-item-padding-right:5px;--awb-item-padding-left:5px;"><div class="awb-toc-el__content"></div></div><div class="fusion-separator fusion-full-width-sep" style="align-self: center;margin-left: auto;margin-right: auto;margin-top:20px;margin-bottom:20px;width:100%;"><div class="fusion-separator-border sep-single sep-solid" style="--awb-height:20px;--awb-amount:20px;--awb-sep-color:var(--awb-color6);border-color:var(--awb-color6);border-top-width:1px;"></div></div><div class="fusion-image-element " style="--awb-margin-top:25px;--awb-margin-bottom:25px;--awb-caption-title-font-family:var(--h2_typography-font-family);--awb-caption-title-font-weight:var(--h2_typography-font-weight);--awb-caption-title-font-style:var(--h2_typography-font-style);--awb-caption-title-size:var(--h2_typography-font-size);--awb-caption-title-transform:var(--h2_typography-text-transform);--awb-caption-title-line-height:var(--h2_typography-line-height);--awb-caption-title-letter-spacing:var(--h2_typography-letter-spacing);--awb-filter:saturate(100%);--awb-filter-transition:filter 0.3s ease;--awb-filter-hover:saturate(0%);"><span class=" fusion-imageframe imageframe-none imageframe-11 hover-type-zoomout"><img decoding="async" width="1536" height="1024" title="blog lvl2" src="https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15.png" alt class="img-responsive wp-image-1687" srcset="https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-200x133.png 200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-400x267.png 400w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-600x400.png 600w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-800x533.png 800w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15-1200x800.png 1200w, https://urbangeoanalytics.com/wp-content/uploads/2025/11/ChatGPT-Image-7-nov.-2025-09_10_15.png 1536w" sizes="(max-width: 640px) 100vw, 400px" /></span></div></div></div></div></div>
<p>The post <a href="https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/">Introducing UVLM: A Free Tool to Compare AI Models That Understand Images</a> appeared first on <a href="https://urbangeoanalytics.com">Urban Geo Analytics</a>.</p>
]]></content:encoded>
					
					<wfw:commentRss>https://urbangeoanalytics.com/introducing-uvlm-free-tool-compare-ai-vision-language-models/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>
