← Back to blog

Published · Updated

Structured Data Extraction: Turning Messy HTML into Clean JSON

Fetching a page is the easy part of web scraping. Turning what you get back into consistent, structured JSON that downstream code can actually rely on is where most of the real engineering work happens. Here's a rundown of the main approaches, roughly in order of how much they depend on a page's specific markup.

1. CSS selectors and tag traversal

The most direct approach: target specific elements by class, tag, or attribute.

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
data = {
    "title": soup.select_one("h1.product-title").get_text(strip=True),
    "price": soup.select_one("span.price").get_text(strip=True),
}

Strength: precise, fast, no dependencies beyond a parser. Weakness: brittle. Any redesign or A/B test that changes class names breaks the selector, and you won't know until the output silently goes empty or wrong.

2. Structured metadata (schema.org, Open Graph, JSON-LD)

Many sites embed structured data directly in the page specifically so search engines and social platforms can parse it reliably, and you can read the same data.

import json
from bs4 import BeautifulSoup

soup = BeautifulSoup(html, "html.parser")
script = soup.find("script", {"type": "application/ld+json"})
data = json.loads(script.string)

Strength: structured by design, far less likely to change than visual markup, often more complete than what's visible on the page. Weakness: not every site includes it, and completeness varies. Some sites only tag a subset of fields.

3. Regex and pattern matching

For simple, consistently-formatted values (phone numbers, prices, dates) embedded in text rather than clean tags, regex can extract a value without needing a full DOM parse.

Strength: works even on malformed or heavily-templated HTML. Weakness: fragile against format variation, and easy to write patterns that silently match the wrong thing.

4. LLM-assisted extraction

A newer approach: feed a chunk of HTML or rendered text to a language model with instructions to extract specific fields as JSON. This handles layout variation well, since the model reasons about meaning rather than matching exact tags. For agentic workflows, keep that extraction distinct from discovery and page reading; our three-tool web-scraping architecture explains why the separation improves routing and context use.

Strength: resilient to markup changes, useful for extracting from genuinely unstructured or highly variable content. Weakness: slower and more expensive per page than selector-based parsing, and needs validation. Models occasionally hallucinate or misparse edge cases, so check the output against an expected schema.

5. Using a source that's already structured

When a managed API covers the data you need, you can consume a documented response instead of maintaining an HTML parser. PrismCrawl provides structured results for search, shopping, maps, travel, reviews, research, and app stores. For a Google Search request, that includes organic listings and supported SERP features. Check the API reference for the response shape of each operation.

Choosing an approach

ApproachReliabilityCoverageCost
CSS selectorsBreaks on redesignAny visible fieldFree (your time to maintain)
Structured metadataStableOnly tagged fieldsFree
RegexFragileSimple patterns onlyFree
LLM-assistedResilientBroadPer-request API cost
Native structured APIVery stableSpecific to that data sourcePer-request API cost

Most production pipelines end up combining two or three of these: structured metadata where available, selectors as a fallback, and an LLM or a dedicated API for the messiest or highest-value data sources.

Frequently asked questions

Why do CSS selector scrapers break so often?

Because they depend on implementation details (class names, DOM structure) that site owners change freely for redesigns, A/B tests, or framework migrations, without any obligation to keep them stable for scrapers.

Is schema.org/JSON-LD data always accurate?

It's generally reliable where present, since sites maintain it for SEO and rich-result purposes, but coverage is inconsistent; always check what fields a given site actually populates before depending on it.

Is there a way to get search data without parsing HTML at all?

Yes. PrismCrawl provides documented JSON responses for search, shopping, maps, travel, reviews, research, and app stores. Your application consumes the data without maintaining HTML selectors for those sources.