How to Scrape Any Website into Markdown (2026 Guide)

Raw HTML wastes up to 96% of your tokens. We ran four real pages through Firecrawl, Jina Reader and two DIY scripts to find the cleanest way to get complete Markdown.

Author
ProxyHorizon Team
Published
September 28, 2026
15 min read
Expert-Verified
How to Scrape Any Website into Markdown (2026 Guide)

Our own guide to deleting incognito history runs about 2,000 words. As raw HTML, it weighs 91,470 tokens.

As Markdown, it’s 3,938.

Same words, same headings, 23 times cheaper to hand to an LLM. That gap is why so many developers now scrape websites into Markdown before an AI model ever reads them.

Here’s the catch nobody mentions. The two tidiest outputs in our test had quietly deleted a third of that article, and their low token counts made the loss look like a win.

So we ran four real pages through five methods and checked what survived. You’ll get working code, the raw numbers, and the one setting that fixes messy docs pages.

TL;DR
  • Markdown cut our four test pages by 70–96% compared with raw HTML.
  • The smallest output isn’t the best one: two tools got there by dropping a whole FAQ.
  • Firecrawl returned every test page complete. Docs sites needed one extra option.
  • Check whether a site already serves Markdown before you scrape it. Some do.

Disclosure: we’re a Firecrawl affiliate, so we may earn a commission if you upgrade through our links. Every tool went through the same tests, and we report the numbers as they came out.

Why Your LLM Wants Markdown, Not HTML

HTML is written for browsers. A typical page wraps its text in scripts, style rules, tracking tags, menus and layout boxes, and a model reads every byte you send it.

That costs you twice. You pay for every token, and the clutter eats context window space the real text needs. In a web scraping pipeline that feeds a search index, the menu text gets indexed too, so searches start matching navigation instead of content.

Markdown keeps only what carries meaning: headings, lists, tables, links and code. Models see it constantly in docs and README files, so they follow its structure well. You can read it too, which matters more than you’d think when something breaks.

Comparison of raw HTML full of scripts, menus and trackers with clean Markdown made of headings, lists and links
Same page, two formats. The model only needs the right-hand side.

The savings are big. Our test product page dropped from 2,123 tokens of HTML to 459 tokens of Markdown. Our Next.js article fell from 91,470 to 3,938. Firecrawl’s homepage claims 93% fewer input tokens, and our four pages landed between 70% and 96%.

In plain English: pay for the page, not the wrapping paper.

Check Whether the Site Already Serves Markdown

Before you scrape anything, spend ten seconds checking whether the site will simply hand you Markdown. More sites do than you’d guess, and none of the guides we read mention it.

Three quick checks cover it:

  • llms.txt. A plain-text index of a site’s pages for AI tools, often linking to Markdown copies. Firecrawl’s own docs publish one.

  • A .md twin. Some docs platforms serve any page as Markdown if you add .md to the URL.

  • An Accept header. Since February 2026, sites on Cloudflare’s paid plans can turn on Markdown for Agents, which converts pages at the edge when a client asks for text/markdown.

Bash
# 1) Is there an llms.txt index?
curl -s https://docs.firecrawl.dev/llms.txt | head -5

# 2) Does this docs page have a .md twin?
curl -sI https://docs.firecrawl.dev/features/scrape.md | grep -i content-type

# 3) Will the server convert the page for you? (GET, keep only the headers)
curl -s -o /dev/null -D - -H "Accept: text/markdown" https://www.cloudflare.com/ \
  | grep -iE "content-type|x-markdown-tokens"

All three worked when we tried them. Cloudflare’s homepage came back as text/markdown with an x-markdown-tokens: 750 header, and one of its blog posts shrank from 315 KB of HTML to 18 KB of Markdown. Use a normal GET for that last check: a HEAD request reported zero tokens, because there was no body to count.

The catch: most sites don’t offer any of this yet. Ours doesn’t, and we checked. When all three come back empty, it’s time to scrape.

The No-Code Way to Convert a Page

Need one page, not a pipeline? Skip the code. Firecrawl’s free website to Markdown converter takes a URL and returns clean Markdown in your browser, with no account and no API key.

Firecrawl’s free website to Markdown converter with a URL box and a Convert to Markdown button, no signup needed
Firecrawl’s free converter, captured September 29, 2026.

Paste the URL, click Convert to Markdown, and copy or download the result. It renders JavaScript first, so it copes with modern single-page apps, not just static HTML.

Best for: one-off pages and quick checks before you automate anything. Not ideal for: hundreds of pages, scheduled jobs or anything that feeds a pipeline. That’s the API’s job.

Scrape Any Page to Markdown with the Firecrawl API

The API does the same job from code, at one credit per page. Grab a free key first. Firecrawl’s docs also allow a few requests without one, but our test machine got an “IP address looks suspicious” refusal, so don’t plan around it.

1Python

Python
# pip install firecrawl-py
from firecrawl import Firecrawl

firecrawl = Firecrawl(api_key="fc-YOUR-API-KEY")

doc = firecrawl.scrape(
    "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
    formats=["markdown"],
)

print(doc.metadata.title)
with open("page.md", "w", encoding="utf-8") as f:
    f.write(doc.markdown)

Books to Scrape is a sandbox built for scraping practice, so it’s a safe first target. Here’s a trimmed look at what came back for that page:

Markdown
- [Home](https://books.toscrape.com/index.html)
- [Books](https://books.toscrape.com/catalogue/category/books_1/index.html)
- [Poetry](https://books.toscrape.com/catalogue/category/books/poetry_23/index.html)
- A Light in the Attic

# A Light in the Attic

£51.77

In stock (22 available)

## Product Information

| UPC | a897fe39b1053632 |
| Product Type | Books |
| Price (excl. tax) | £51.77 |

Notice two quirks. The breadcrumb links survived the main-content filter, and the product table has no separator row under its first line, because the page’s table has no header section. A model reads it fine, but a Markdown renderer won’t draw it as a table.

2Node.js

JavaScript
// npm install firecrawl
import { Firecrawl } from 'firecrawl';

const firecrawl = new Firecrawl({ apiKey: 'fc-YOUR-API-KEY' });

const doc = await firecrawl.scrape(
  'https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html',
  { formats: ['markdown'] },
);

console.log(doc.markdown);

3cURL

Bash
curl -X POST https://api.firecrawl.dev/v2/scrape \
  -H 'Content-Type: application/json' \
  -H 'Authorization: Bearer fc-YOUR-API-KEY' \
  -d '{"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html", "formats": ["markdown"]}'

The response puts the page under data.markdown, with the title, final URL and status code in data.metadata. Markdown is the default format, so you can leave formats out entirely.

We Ran 4 Real Pages Through 5 Methods

Every guide says Markdown saves tokens. None that we found showed what each method keeps and what it throws away, so we tested it ourselves in September 2026.

We picked four pages that trip up converters in different ways:

  • A static product page on Books to Scrape.

  • A Quotes to Scrape page where JavaScript injects every quote.

  • Python’s documentation page for html.parser.

  • Our own Next.js article on deleting incognito history.

Each page then went through the five methods in the table below, with the raw HTML as the baseline. Both DIY runs worked on the HTML a plain HTTP request returns, just like a requests script: markdownify converted all of it, while trafilatura pulled out the main text first.

Jina Reader ran on its free endpoint. Firecrawl ran through its hosted MCP server, which calls the same scrape endpoint as the API, on default extraction settings with the cache switched off. We counted tokens with OpenAI’s o200k_base tokenizer, the one GPT-4o uses.

23×
Smaller than raw HTML
Our Next.js article, Firecrawl Markdown
0 of 10
Quotes the DIY scripts found
JavaScript page, markdownify and trafilatura
9
FAQ answers two tools dropped
Jina Reader and trafilatura
−33%
Tokens after one setting
Docs page with includeTags

Source: ProxyHorizon tests, September 2026. Tokens counted with OpenAI’s o200k_base tokenizer.

PageRaw HTMLmarkdownifytrafilaturaJina ReaderFirecrawl
Product page (static)2,123467357488459
Quotes page (JavaScript)1,48466, no quotes10, no quotes268367
Python docs page15,6114,1603,0623,4524,688 (3,119 tuned)
Our Next.js article91,4705,0671,801, FAQ missing1,892, FAQ missing3,938

Read the table with one question in mind: which small numbers came from removing noise, and which came from losing content? The next three sections answer it.

1JavaScript Pages Break the DIY Route

The quotes page ships almost empty HTML. All ten quotes sit inside a script and appear only after the browser runs it, the same pattern many modern web apps use.

So markdownify returned 66 tokens and trafilatura returned 10, with zero quotes between them. Both parse the HTML you hand them and neither runs JavaScript, which is why DIY scrapers usually end up bolting on a headless browser.

Firecrawl and Jina Reader both render the page first, and both found the quotes. Firecrawl kept all ten with their authors and tags, though the tags ran together into strings like “changedeep-thoughtsthinkingworld”. Jina Reader dropped every tag and lost the first quote’s author.

2The FAQ That Quietly Vanished

On our own article, Jina Reader and trafilatura produced the smallest Markdown, about 1,800 to 1,900 tokens against Firecrawl’s 3,938. It looks like a clear win until you open the output.

Both kept the “Frequently Asked Questions” heading and dropped all nine answers under it. That’s 810 words, roughly a third of the page’s text.

FAQ section of our test article with questions in collapsed accordion rows, the answers Jina Reader and trafilatura dropped
The FAQ on our test article, captured September 29, 2026. Each answer sits inside a collapsed row like the highlighted one.

The answers are right there in the HTML, inside collapsed accordion panels whose questions are buttons. Extractors that score blocks for “article-ness” tend to treat widgets like that as page furniture. Firecrawl kept all nine answers, plus some genuine noise: the byline, a related-posts block and a table-of-contents stub.

Our take: a few lines of noise cost you fractions of a cent. A missing third of the page costs you wrong answers, and nothing warns you it happened. Check what your converter dropped, not just how small the output got.

3Docs Sites Need One Extra Setting

Firecrawl’s weakest result came on the Python docs. Its main-content filter kept both navigation bars, the language and version switchers, and ten “Copy” button labels, which made its output the biggest of the four Markdown methods.

The fix took one look at the page source. Sphinx, the tool behind Python’s docs, wraps the real content in div.body, so we told Firecrawl to keep only that:

Python
doc = firecrawl.scrape(
    "https://docs.python.org/3/library/html.parser.html",
    formats=["markdown"],
    include_tags=["div.body"],                               # the docs' content wrapper
    exclude_tags=["button", ".copybutton", "a.headerlink"],  # copy buttons and ¶ links
)

Output fell from 4,688 to 3,119 tokens, a third smaller, with all eleven code examples intact. Most docs generators use a similar wrapper, so find it once per site and reuse it.

The Options That Clean Up Your Markdown

You’ll rarely need more than two or three of these, but it pays to know they exist. Defaults come from Firecrawl’s API reference, and the Python SDK uses the same names in snake_case.

OptionDefaultUse it when
formatsMarkdownYou also want HTML, links or a screenshot
onlyMainContenttrueSet it to false only if you want menus and footers
includeTagsNoneThe page has a clear content wrapper, like main, article or div.body
excludeTagsNoneWidgets slip through, like related posts, share bars or author boxes
waitFor0 ms, on top of smart waitingContent shows up a second or two after the page loads
actionsNoneYou need to scroll, click “Load more” or type before scraping
maxAgeUp to 2 daysPrices, stock or news: set 0 to force a fresh fetch
parsersPDF parsing onYou’re converting PDFs
blockAdstrueLeave it on. It blocks cookie pop-ups too
location, mobileUS, desktopThe page changes by country or device

Both tag filters take CSS selectors and run against the original page, so anything you can target in your browser’s DevTools works.

How to Handle Infinite Scroll, Logins, PDFs and Bot Walls

“Any website” includes the awkward ones. We didn’t test these features ourselves, so this section leans on Firecrawl’s docs and the SDK’s own type definitions, which we checked every sample against.

Infinite scroll and “Load more” buttons. Browser actions run before the scrape. You can chain up to 50 of them, as long as the waits add up to 60 seconds or less:

Python
doc = firecrawl.scrape(
    "https://quotes.toscrape.com/scroll",
    formats=["markdown"],
    actions=[
        {"type": "scroll", "direction": "down"},
        {"type": "wait", "milliseconds": 1500},
        {"type": "scroll", "direction": "down"},
        {"type": "wait", "milliseconds": 1500},
        # a "Load more" button works too:
        # {"type": "click", "selector": "button.load-more"},
    ],
)

PDFs. Pass a PDF’s URL like any other page. PDF parsing is on by default and billed per PDF page, so a long report costs more than a single web page.

Logins. For pages you’re allowed to see, like your own dashboard, pass your session cookie in headers. Firecrawl doesn’t cache requests that carry custom headers, so the private page isn’t stored for anyone else. Read the site’s terms before you automate a logged-in area.

Bot walls. The proxy option defaults to auto, which switches to enhanced proxies when a basic request fails, at no extra credit cost. If you run your own scrapers instead, you’ll need residential proxies and a plan for getting past Cloudflare.

Turn a Whole Website into Markdown Files

For a whole docs site or blog, don’t scrape URL by URL. Map the site to see what’s there, then crawl only the part you need. Firecrawl’s docs price a map call at one credit however many URLs it returns, and a crawl costs one credit per page.

Python
from pathlib import Path
from urllib.parse import urlparse
from firecrawl import Firecrawl

firecrawl = Firecrawl(api_key="fc-YOUR-API-KEY")

# 1) See what's there first: one map call costs one credit
site = firecrawl.map("https://docs.firecrawl.dev", limit=500)
print(len(site.links), "URLs found")

# 2) Crawl only the section you need, with a hard page cap
job = firecrawl.crawl(
    "https://docs.firecrawl.dev",
    include_paths=["^/features/.*"],
    limit=50,
    formats=["markdown"],
    only_main_content=True,
)
print(job.status, job.completed, "pages,", job.credits_used, "credits")

# 3) Save every page as its own .md file
out = Path("markdown")
out.mkdir(exist_ok=True)
for page in job.data:
    url = page.metadata.source_url or page.metadata.url or ""
    name = urlparse(url).path.strip("/").replace("/", "_") or "index"
    (out / f"{name}.md").write_text(page.markdown or "", encoding="utf-8")

Two settings protect your credits. include_paths takes regular expressions matched against each URL’s path, and limit caps the page count. Leave the limit out and a crawl can run to 10,000 pages, which is ten months of the free plan.

The crawler also respects robots.txt unless you set ignore_robots_txt=True, which you shouldn’t. For queues, webhooks and bigger jobs, see our guide to scraping large websites with Firecrawl.

Chunk Your Markdown for RAG Without Breaking Code

If the Markdown feeds a RAG system, split it at headings rather than at a fixed character count. Heading-based chunks keep each idea together, and the heading path makes a handy label for search results.

The trap is code. A Python comment starts with #, exactly like a Markdown heading, so naive splitters slice code blocks in half. This function tracks code fences to avoid that:

Python
import re

HEADING = re.compile(r"^(#{1,3}) +(.+)")
LINK = re.compile(r"\[([^\]]+)\]\([^)]*\)")


def chunk_markdown(markdown, url, max_chars=2000):
    """Split Markdown into heading-aware chunks. Never cuts a code block in half."""
    chunks, path, block, in_code = [], [], [], False

    def flush():
        text = "\n".join(block).strip()
        if text:
            section = " > ".join(title for _, title in path)
            chunks.append({"url": url, "section": section, "text": text})
        block.clear()

    for line in markdown.splitlines():
        if line.startswith("```"):
            in_code = not in_code
        heading = None if in_code else HEADING.match(line)
        if heading:
            flush()
            level = len(heading.group(1))
            while path and path[-1][0] >= level:
                path.pop()
            path.append((level, LINK.sub(r"\1", heading.group(2)).strip()))
        elif not in_code and not line.strip() and sum(map(len, block)) > max_chars:
            flush()
        block.append(line)
    flush()
    return chunks


# Usage with the crawl results from the previous section
for page in job.data:
    for chunk in chunk_markdown(page.markdown or "", page.metadata.source_url):
        print(chunk["section"], len(chunk["text"]))

We ran it over our test outputs. Our incognito article became 29 chunks and the Python docs page became 7, with no broken code blocks and no headings pulled from inside code. Store the url and section with each chunk so answers can cite their source. The full pipeline is in our guide to using Firecrawl for RAG.

What It Costs, and What You Save

Firecrawl bills in credits: one per basic page scraped, crawled or mapped. The free plan’s 1,000 credits refresh every month and don’t need a card. These prices took effect on September 4, 2026.

PlanMonthlyBilled yearlyCredits a monthConcurrent requests
Free$0–1,0002
Hobby$19$16/mo5,0005
Standard$99$83/mo100,00025
Growth$399$333/mo500,00050
Scale$749$599/mo1,000,000100
Firecrawl pricing page: 1 credit per page, 1,000 free credits a month with no card, and 403 or 404 pages still cost 1 credit
Firecrawl’s pricing page, captured September 29, 2026 (USD, effective September 4, 2026).

Mind the fine print highlighted there. A page that answers with a 403 or 404 still costs a credit, so clean your URL list before a big run.

A worked example. Say you convert 10,000 pages a month. Hobby includes 5,000 credits, and pay-as-you-go adds 1,000 more per $5, so those 10,000 pages cost about $44. Past roughly 20,000 pages a month, Standard at $99 works out cheaper.

Now the other side of the ledger. If your pages weigh as much as our Next.js article, Markdown saves about 87,500 tokens each, or 875 million across those 10,000 pages. At $1 per million input tokens, a round number for the math rather than a real quote, that’s $875 you don’t spend on the model.

Our take: on heavy modern pages, the conversion pays for itself many times over. On light static pages like our product page, the same math saves about $17 per 10,000 pages, and a free DIY script may be all you need. Every plan detail is in our Firecrawl pricing breakdown.

Picking Between Firecrawl, Jina Reader and DIY Scripts

Here’s how the main options compare. We tested four of them, and the Crawl4AI and MarkItDown rows come from their official docs.

MethodJavaScriptBest forWatch out for
FirecrawlYesAny page or whole site; 1,000 free credits a monthSome widgets slip through; docs sites need includeTags
Jina ReaderYesOne-off articles; 20 free requests a minuteDropped our FAQ and a quote’s author
Crawl4AIYesSelf-hosted, open-source crawlsYou run the browsers and proxies
MarkItDownNoPDFs, Word and other filesKeeps menus and footers on web pages
markdown­ify (DIY)NoStatic HTML you already haveKept menus and footers; blind to JavaScript
trafilatura (DIY)NoStatic articles and blogsDropped our FAQ; blind to JavaScript
Native MarkdownSite does itSites with llms.txt, .md pages or the Accept headerStill rare

Firecrawl for Markdown: The Short Version

Pros5
  • Rendered every JavaScript page we tested
  • Keeps links absolute, so citations still work
  • Scrape, map and crawl from one API
  • 1,000 free credits a month, no card
  • Python and Node SDKs, plus a CLI and MCP server
Cons4
  • Some widgets survive the main-content filter
  • Docs sites need includeTags for clean output
  • 403 and 404 pages still cost a credit
  • The self-hosted version lacks the anti-bot engine

Verdict

The best default for turning real pages into complete Markdown. Add includeTags where a site needs it and it’s hard to beat.

Our take: Firecrawl is the best default. It was the only method that returned all four test pages complete, from the JavaScript quotes to the collapsed FAQ.

Jina Reader is a fine free choice for one-off articles. Crawl4AI suits teams who want open source and can run their own browsers and proxies. For the wider field, see our Firecrawl alternatives roundup.

Mistakes That Waste Tokens, Credits or Content

1Treating Scraped Text as Trusted Input

Scraped Markdown goes straight into a model’s context, so anything on the page can talk to your model. Pages increasingly carry text aimed at AI agents. Several Firecrawl pages we pulled for this article included notes addressed to AI agents, and so did the API’s error message.

Those notes were harmless. The same channel carries prompt injection. Wrap scraped content in clear delimiters, tell the model it’s data rather than instructions, and never let it trigger tools on its own.

2Letting the Cache Serve Stale Data

Firecrawl can answer from its cache to save time. On the raw API, the default accepts a copy up to two days old, and cached results still cost a credit. That’s fine for docs. For prices, stock levels or news, pass max_age=0 to force a fresh fetch.

3Throwing Away the Source URL

Merging a whole crawl into one big file is tempting. Do that and you lose what your RAG answers need most: where each fact came from. Keep one file per page, or at least store the URL and fetch date with every chunk, so you can cite sources and refresh stale pages later.

4Forgetting Whose Content It Is

Converting a page doesn’t change who owns it. Respect robots.txt and the site’s terms, keep request rates polite, and be careful with personal data, which laws like the GDPR protect however you collected it. Summarizing public pages for your own use is one thing. Republishing them wholesale is another.

Frequently Asked Questions

Yes, in several ways. Firecrawl’s website to Markdown converter works in your browser with no signup, and its free API plan includes 1,000 credits a month, enough for about 1,000 pages. Jina Reader allows 20 requests a minute without a key. Python libraries like markdownify cost nothing, but they can’t run JavaScript, so they come back empty on many modern sites.
Map the site first, then crawl only the paths you need. With Firecrawl, one map call costs a credit and returns the site’s URLs, and a crawl returns every page as Markdown at one credit per page. Set a limit and include_paths so a big site can’t drain your credits, then save each page as its own .md file.
It renders each page in a real browser before converting it, so content that JavaScript adds ends up in the Markdown. In our test it captured all ten quotes on a JavaScript-only page where two Python scripts found none. For content that loads late or hides behind a button, add waitFor or browser actions like scroll and click.
For straight conversion, markdownify and html2text are the usual picks. Trafilatura is better when you want the main article first, and it outputs Markdown directly. None of them run JavaScript, and in our test trafilatura also dropped a whole FAQ section. Use them on static pages you understand, and a rendering API for everything else.
Usually. Markdown keeps headings, lists, tables and links, which help a model see how ideas relate and help a RAG system split pages at sensible points. Plain text flattens all of that for a tiny saving, since Markdown syntax adds only a few characters per heading or link. Raw HTML is the format to avoid.
Only pages you’re allowed to access, such as your own account or an internal tool. Firecrawl can send your session cookie through its headers option, and it won’t cache requests that carry custom headers. Check the site’s terms first: many ban automated access to logged-in areas, and collecting other people’s data can break privacy law.
Scraping public pages is legal in many countries, but it depends on the site’s terms, copyright and any personal data involved. Respect robots.txt, which Firecrawl’s crawler follows by default, keep your request rate polite, and don’t republish what you collect as your own work. For commercial projects involving personal data, get legal advice for your country.

The Bottom Line

Turning a website into Markdown is easy. Getting Markdown that’s both small and complete takes a little care, and that’s the part most guides skip.

Check for a native Markdown version first. When there isn’t one, Firecrawl is the default we’d reach for. It rendered every page we tested, kept content the tidier tools dropped, and scales from one URL to a whole site.

Add include_tags on docs sites, and read what your converter left out before you trust its token count.

Start on the free plan with the three messiest pages you actually care about. You’ll know within ten minutes whether it fits.