Open-weight models • Runs on your hardware • RAG • Reproducible by design • Fine-tuning

UNDERSTANDING THE TECHNOLOGY

Language Models That Run Where Your Data Lives

Large language models (LLMs) are usually consumed as online services: a question and the documents attached to it are sent to a provider, processed on its servers, and the answer comes back. For many organisations this is not acceptable. Research data held under an ethics protocol, legal files, medical records, unpublished manuscripts, tenders and internal strategy documents cannot leave the building.

Open-weight models change this. Families such as Qwen, Llama, Gemma and Mistral publish their trained weights, which means the model itself can be downloaded and run on a workstation, on a laptop with a graphics card, or on a server inside your own network. Once installed, it needs no internet connection. Nothing is sent anywhere, and nothing changes unless you decide to change it.

The same models can do more than answer. Connected to tools (a document search, a file reader, a database, a script), a local model becomes an agent that carries out a task in several steps, on your files and inside your network.

  • Confidentiality. Documents, questions and answers stay on your hardware.
  • Reproducibility. A fixed model, fixed settings and a fixed seed return the same answer today and in two years. A hosted model can be updated or withdrawn without notice.
  • Control. The model, its instructions, its sources and its safety behaviour are chosen for your use case, not inherited from a general-purpose product.
  • Predictable cost. No fee per question and no subscription: the cost is the hardware you already own or choose to buy.
  • Independence. No account, no vendor lock-in, no exposure to changes in a provider’s terms.

OUR APPROACH

How We Build Tailored Local LLMs and Agents

Urban Geo Analytics (UGA) designs, evaluates and delivers language model and agent systems that run entirely on the client’s infrastructure. Every system is built for one organisation and one purpose: the model, the knowledge it draws on, the way it answers and the interface it is used through are all part of the specification.

Model Selection & Hardware Sizing

There is no best model, only a best model for a task and a machine. We select among open-weight families and sizes by testing candidates on the client’s own questions and documents, then fit the chosen model to the available hardware. Quantisation (4-bit and GGUF formats) lets an 8-billion-parameter model run on a consumer graphics card with 8 GB of memory; larger cards and servers open the way to larger models and longer contexts. We measure memory use, speed and answer quality for each configuration and deliver the figures, so that the hardware decision rests on evidence. We also check the licence attached to each model, since open weights are not always free for commercial use.

Retrieval-Augmented Generation on Your Documents

A language model on its own answers from what it learned during training. Retrieval-augmented generation (RAG) connects it to your documents instead: reports, notes, spreadsheets and PDFs are indexed locally, the passages relevant to each question are retrieved, and the model answers from them, showing the passages it used so that every statement can be checked. We configure how strictly the model must keep to the documents, from “documents only, and say so when the answer is not there” to documents completed by the model’s general knowledge. RAG is optional: some projects need a plain local model with a well-designed system prompt and nothing else.

Local Agents: From Answering to Doing

An agent is a language model connected to tools: functions it can call, files it can read, scripts it can run. Instead of answering one question, it pursues a task through several steps: it plans, calls a tool, reads the result and decides what to do next. Run locally, an agent works on your files and systems without exposing any of them to an outside service. We build each agent around a defined tool set, typically document search, file and spreadsheet reading, database queries, Python scripts and, in our own field, geospatial tools. Typical tasks are extracting the same fields from every document of a collection into a table, compiling a standard report from a set of files, or checking documents against a list of rules.

Compact open-weight models are less dependable than large hosted ones over long chains of actions, and we design for it: short, well-specified steps, structured tool calls that are validated before they are executed, and a test set that measures how often the task is completed correctly before the agent is put to work.

Permissions, Approval & Audit

What an agent may do is decided before it runs. Each tool is granted explicitly, with its scope: which folders, read-only or read-write, which databases. Actions that modify or delete something can be made to wait for human approval. Every step, tool call and result is logged, so that a run can be inspected afterwards and replayed.

Fine-Tuning, Compression & Customisation

When instructions and retrieval are not enough, the model itself can be adapted. We fine-tune open-weight models on the client’s data with parameter-efficient methods (LoRA, QLoRA) to teach a vocabulary, a writing style, a classification scheme or a structured output format. We reduce models through quantisation and pruning so that they fit smaller hardware or answer faster, and we measure what each reduction costs in quality. Lighter customisation, such as a persona, a tone, a fixed output template or a controlled reasoning budget, is handled through the prompt architecture and needs no training.

Safety Behaviour as a Project Parameter

General-purpose assistants apply one moderation policy to everyone. With a local model, that policy is part of the specification. A public-facing assistant can be given strict guardrails and a documented refusal policy. A research team working on sensitive material (health, law, security, or the study of harmful content itself) may need a model that analyses its sources without refusing. We define this with the client, document the choice, and deliver the system for use within the client’s legal and ethical framework.

Evaluation & Reproducibility

We test before we deliver. Candidate models are evaluated on the client’s tasks for accuracy, for faithfulness to the sources (does each statement in the answer follow from the retrieved passages?) and for stability across repeated runs; agents are evaluated on how often they complete the task correctly. Decoding settings, seeds, prompts and model versions are recorded with every result, so that an answer can be traced and reproduced. This is the discipline we apply in our own research on the faithfulness of open-weight models.

APPLICATIONS

Who Uses Local LLMs and Agents

Research teams use local models to query interview transcripts, archives and literature collections under data protection constraints, and to make the language-model step of a study reproducible and citable. Law firms, health organisations and public bodies use them to search and summarise files that are not allowed onto external servers. Companies use them to give staff an assistant over internal documentation, procedures and technical manuals without exposing intellectual property, and agents to turn piles of contracts, reports or forms into structured tables. NGOs and field organisations use them where connectivity is poor or the subject matter is sensitive. In our own domain, local models are the reasoning layer of spatial tools: conversational access to geospatial data and automated urban diagnostics (see Geospatial Analysis & Urban Intelligence).

DELIVERY

From Notebook to Installable Application

The same system can be delivered in several forms. For researchers and data teams, a documented Jupyter notebook exposes every setting: model, retrieval, persona, temperature, seed and reasoning budget. For non-technical users, we package the system as a Windows desktop application that installs like any program, downloads the model once at first launch, and then works with no internet connection: a chat window with cited sources, a document library, saved and searchable conversations, and settings remembered between sessions. For organisations, we deploy the model on an on-premise server behind a private API, with authentication and logging, so that a whole team shares one installation. Where independence from a single cloud provider matters more than physical locality, the same models can be hosted on decentralised compute (see IoT, Blockchain & Decentralised AI).

OUR TOOLS

RawRAG and Our Research on Open-Weight Models

RawRAG is our desktop application for question answering over private documents. It runs the Qwen3-8B model locally, measures the available graphics memory before loading to choose the best configuration for the machine, falls back to the processor when no graphics card is present, and cites the passages behind every answer. It is the reference implementation of the desktop delivery described above.

Our research on open-weight language models includes NSCR-LLM (Network-based Spatial Context Retrieval for LLMs), an open pipeline and benchmark that measures, claim by claim, how faithfully these models reason from the spatial context they are given, across model families, sizes and repeated runs.

WORKING WITH US

Engagement Options

A feasibility study tests two or three candidate models on a sample of your documents and questions and delivers a recommendation with measured quality, speed and hardware requirements. A turnkey assistant is a complete local system (model, document index, interface) installed on your machines, with a user guide and a training session. A custom model engagement covers fine-tuning or compression for a specific task, with evaluation before and after. A local agent engagement defines the task, the tools and the permissions, and delivers the agent with its test results and its logging. A research partnership integrates a reproducible local LLM component into a funded project, with the documentation needed for publication. In every case you receive the system, its configuration and its documentation, and you keep full ownership of your data.

Need a language model or an agent that stays on your machines? Get in touch.