In the first post in this series, we made a simple point: AI is pattern recognition, not magic. The same technology that puts a top hat on your nephew's cat can classify millions of points in a LiDAR dataset. One is a party trick. The other is a production tool.
"AI" isn't one thing. It's a whole zoo of different model types, each built for a different kind of task, each with its own strengths and blind spots. The difference between picking the right model for the job and picking the wrong one is the difference between a tool that saves your team hours and a toy that wastes them.
Most people in this industry, and most people in general, hear "AI" and think of ChatGPT. That's like hearing "surveying instrument" and thinking only of a total station. You wouldn't use a total station to process aerial imagery. You wouldn't use a drone to check a benchmark elevation. Different tools for different jobs. AI works the same way.
This post is your field guide to the AI model zoo. By the end, you'll know enough to evaluate vendor claims, ask better questions about the tools you're being pitched, and understand why "we use AI" is about as informative as "we use instruments."
Large Language Models: The Colleague Who's Read Everything But Never Set a Monument
Large language models, or LLMs, are the ones you've probably already used. ChatGPT, Claude, Gemini, Meta's Llama, Mistral. These models were trained on massive amounts of text, billions of pages of books, articles, websites, code, and documents, and they learned to predict what words should come next in a sequence based on the patterns they found.
LLMs can draft legal descriptions, summarize 40-page title commitments, translate dense regulatory language into plain English, generate report narratives from structured data, and write project proposals. They process text and produce text, faster than any human can type.
An LLM is that colleague who's read every textbook, every state statute, every ALTA standard, every deed ever recorded. They can talk about any of it fluently and at length. Impressive recall. Articulate delivery.
They've also never held a prism pole. Never walked a property line. Never stood in a county recorder's office arguing about a rejected plat. Their knowledge is broad but purely textual. They know what words appear near other words. They don't understand what those words mean in the physical world.
LLMs will also confidently produce output that sounds authoritative and is factually wrong. Ask one to generate a metes-and-bounds description, and it will hand you something that reads like a perfectly formatted legal description. The bearings might not close. The distances might contradict the deed. It doesn't check its own math, because it isn't doing math. It's predicting which text pattern should come next.
Where LLMs earn their keep: Drafting text-heavy deliverables. Summarizing research packages. Writing proposals and client emails. First drafts that a professional refines. Business tasks where speed matters more than spatial precision.
Where they'll burn you: Anything involving coordinates, calculations, spatial relationships, or factual claims that need verification. A starting point, never a final answer, for work that carries professional liability.
Computer Vision Models: The Eye That Never Blinks
Computer vision models process images and visual data instead of text. They were trained on millions of labeled images and learned to recognize patterns in pixels the way LLMs recognize patterns in words.
A computer vision model is a technician with extraordinary eyesight and limitless patience. It can spot aerial targets in flight imagery at a glance, across thousands of photos, without fatigue. It can read a scanned plat, including inferring and interpolating information where ink has faded or text has blurred from decades of copying. It can classify every pixel in a satellite image as building, road, vegetation, water, or bare earth.
Computer vision encompasses several specialized capabilities. Image classification assigns a label to an entire image: "residential parcel," "commercial site," "wetland." Object detection finds specific things within an image and draws a box around each one: every manhole cover in a flight corridor, every utility pole along a route. Segmentation goes further, classifying every individual pixel as road, building, vegetation, or water, and separating individual objects of the same type into distinct entities.
For geospatial professionals, computer vision is one of the most production-ready applications of AI in the industry today. ESRI has built over 100 pretrained deep learning models into ArcGIS for tasks like feature extraction, land cover classification, and change detection. These aren't experimental. Firms are using them on real projects, processing real data, and producing real deliverables.
Where computer vision earns its keep: Aerial and satellite image analysis. Point cloud classification. Feature extraction. Change detection between survey epochs. Reading and digitizing scanned documents.
Where it falls short: Computer vision sees pixels. It doesn't read legal descriptions or understand property law. It doesn't know the building it detected is encroaching on an easement. Professional interpretation is still required.
Vision-Language Models: When Seeing and Reading Combine
Vision-language models, or VLMs, are a newer category that merges computer vision with language processing. Models like GPT-4V (OpenAI), Gemini (Google), and Claude (Anthropic) can accept both images and text as input. You can show them a photograph, a scanned document, or a map, and ask questions about it in plain English.
This is where things get particularly interesting for surveying and geospatial work. A VLM can look at a scanned subdivision plat and answer questions like "What are the lot dimensions for Lot 14?" or "Does the dedication language reference any easements?" It combines the image processing capability of computer vision with the language comprehension of an LLM.
Feed a VLM a stack of scanned historical plats, prior surveys, and deed exhibits, then ask it to extract specific information, flag inconsistencies, or summarize what it finds. It reads the visual document the way a human would, except it doesn't get tired after the fifteenth page and processes the stack in minutes instead of hours.
VLMs are still maturing. They can misread degraded text, hallucinate details that aren't on the page, and struggle with complex tabular data in low-quality scans. But as a first pass on document review, as a way to triage a 200-page title commitment package before a licensed professional digs in, they're already saving real time.
Where VLMs earn their keep: Scanned document interpretation. Visual Q&A on maps, plats, and imagery. Extracting structured data from unstructured visual sources.
Where they fall short: Accuracy on degraded inputs is inconsistent. Professional review is mandatory.
Foundation Models and Geospatial-Specific AI: Built for Your Data
Foundation models are large-scale AI models trained on broad datasets so they can be adapted to many specific tasks. All LLMs are foundation models. But not all foundation models are LLMs. Some are trained on images, some on scientific data, some on Earth observation data. That last category is where the geospatial industry should be paying close attention.
IBM and NASA released Prithvi-EO-2.0, a geospatial foundation model with 600 million parameters, six times larger than the first version, developed by 42 researchers across 12 institutions. IBM has since released dramatically smaller versions that can run on a smartphone while maintaining performance. The Clay Foundation Model, another open-source project, was built specifically for Earth observation data. Point cloud-specific architectures like PointNet++, RandLA-Net, and KPConv were designed to process three-dimensional spatial data natively.
A general-purpose AI model trained on internet text and stock photos doesn't understand coordinate systems, projections, spatial relationships, or the structure of LiDAR data. A geospatial foundation model does. It was trained on satellite imagery, elevation data, and remote sensing inputs. A general-purpose AI is a bright college graduate who's never worked in the field. A geospatial foundation model is the senior surveyor who's worked every terrain type across multiple states and adapts to new regions fast because the fundamentals are already second nature.
Where geospatial foundation models earn their keep: Land cover classification. Change detection from satellite and aerial imagery. Environmental monitoring. Terrain analysis. Any task where understanding the spatial structure of the data is the whole point.
Where they fall short: These models are powerful but still require professional interpretation. They don't replace the licensed professional who understands what the classification means for a boundary determination, a construction project, or a regulatory compliance question.
The Real Power: Models Working Together
This is the part most people miss entirely.
You don't pick one model and use it for everything. You build workflows where different models handle different steps, each one doing the thing it's best at, and the whole pipeline produces a result that no single model could achieve alone.
A document processing workflow might run like this: a computer vision model reads a scanned plat and extracts line work, lot dimensions, and text annotations. An LLM interprets the legal description, cross-references it against deed language, and flags inconsistencies. A second LLM reviews the first model's output for errors, acting as an automated QC pass. A VLM checks the original scan against the extracted data to confirm nothing was missed or misread.
Four models. One workflow. Each one doing what it was built for.
This is how Enspectri approaches workflow automation for surveying and geospatial teams. Not one model trying to do everything, but purpose-built pipelines where the right model handles the right task, with professional review built into the process. The same principle behind how your firm already works: field crew collects, technician processes, project manager reviews, principal signs. Different expertise at each stage. AI workflows follow the same logic.
You can also use one model to check another. Run the same extraction through two different models and compare results. Where they agree, confidence is high. Where they disagree, a human reviews. Quality control built into the pipeline itself.
Workflows don't have to stay within a single project task. Run one set of models to produce a deliverable, then run a completely separate set to QA it, the same way you'd have one surveyor prepare a plat and a different one review it. AI doesn't eliminate checking. It makes checking faster, more consistent, and more thorough.
Quick Reference: Which Model for Which Job
Text in, text out? LLM. Report drafting, legal description interpretation, proposals, regulatory research. Fast and fluent as a first draft. Always verify.
Images in, analysis out? Computer vision. Aerial imagery, point cloud classification, feature extraction, change detection. Mature and production-ready.
Images in, answers out (in plain language)? Vision-language model. Document Q&A, visual inspection, extracting information from scanned records.
Geospatial data in, geospatial intelligence out? Foundation models trained on Earth observation or spatial data. Land cover classification, environmental analysis, terrain modeling.
Multiple steps, multiple data types? A multi-model pipeline. The most valuable AI implementations chain several models together, with professional review at critical decision points.
When a vendor tells you "we use AI," your follow-up should be: which kind, and how do they work together? The answer tells you whether they've built something for your workflow or bolted a chatbot onto a sales page.
What This Means for You
Only 27% of AEC professionals currently use AI in their work, according to a 2025 Bluebeam survey of over 1,000 industry professionals. Among those who do, 68% have already saved at least $50,000 and nearly half have saved 500 to 1,000 hours. 94% plan to increase their AI usage in 2026.
The firms pulling ahead aren't the ones who tried ChatGPT once and decided AI was either too simple or too complicated. They figured out that "AI" is a toolkit, not a single tool, and the value comes from matching the right model to the right task within a workflow designed by people who understand the work.
AI isn't a black box. It's a zoo. Once you learn which animal does what, the mystery evaporates and the practical value comes into focus.
The labor pressures aren't easing up. The firms that thrive with the teams they have will treat AI the way they treat every other professional tool: understand what it does, learn its limitations, put it to work where it earns its keep.
Next in this series: a deep dive on large language models. How they work under the hood, who's building the major ones, and how surveying professionals are using them in production today.
This is Part 2 of the AI for Surveyors and Geospatial Professionals blog series by Enspectri. Built by industry insiders who've run survey crews, scaled geospatial operations, and shipped production software. See how Enspectri's AI-powered tools deliver real results for real workflows.
Sources
- Bluebeam, "Building the Future: AEC Technology Outlook 2026" (October 2025) (27% adoption, $50K+ savings, 500-1,000 hours saved, 94% plan to increase usage)
- ASCE, "Architecture, Engineering, Construction Sector Slow to Adopt AI, Survey Shows" (December 2025) (AEC adoption context)
- ESRI, "Geospatial Artificial Intelligence" (100+ pretrained deep learning models in ArcGIS)
- IBM Research, "Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications" (600M parameter geospatial foundation model, 42 researchers, 12 institutions)
- AWS, "What Are Foundation Models?" (foundation model definition and taxonomy)
- NVIDIA, "What Are Vision-Language Models?" (VLM definition and capabilities)
- IBM, "What Are Vision Language Models (VLMs)?" (VLM technical overview)
