How to Scrape Websites with n8n (Legally and Reliably)
This page may contain affiliate links.
Scraping gets a bad reputation, partly because a lot of tutorials show you how to hammer a site with 50 concurrent requests and then act surprised when Cloudflare blocks you. But scraping is genuinely useful. Maybe you want to track UK house prices on Rightmove, monitor competitor pricing on your own ecommerce store, or pull job listings from a niche board into a Notion database. n8n is a brilliant tool for this because it handles scheduling, retries, and the awkward glue between HTTP requests and databases without you writing a full app.
This guide covers the legal side first (because it matters and most articles skip it), then walks through a reliable n8n scraping pattern you can actually run in production.
Know the legal and ethical ground rules first
In the UK, the key legal touchpoints are the Computer Misuse Act 1990, the GDPR (if you're scraping personal data), and the terms of service of the site itself. Breaking a site's ToS isn't automatically a criminal offence, but it can be a breach of contract, and if you ignore a cease and desist you're inviting trouble.
The practical rules I follow:
- Check robots.txt first. Visit
https://example.com/robots.txt. If a path is disallowed, don't scrape it. It's not legally binding, but it's the clearest signal of intent a site owner can give you. - Avoid personal data unless you have a lawful basis. Scraping names, emails, or profiles is a GDPR issue the moment you store them. Scraping prices and product descriptions usually isn't.
- Respect rate limits. One request every 2–5 seconds is polite. If a site publishes an API, use it instead.
- Don't republish copyrighted content wholesale. Using prices as data points is fine; copying 5,000 words of editorial is not.
- Identify yourself. Set a User-Agent that names your project and, ideally, a contact URL.
If you're scraping at any scale for a business, get proper legal advice. I'm a blogger, not your solicitor.
Building your first n8n scraper: the reliable pattern
The naive approach is a single HTTP Request node feeding an HTML Extract node. It works until it doesn't. Here's the pattern I actually use:
1. Schedule Trigger. Start with a Schedule Trigger node. For most price-tracking jobs, once or twice a day is plenty. Running every 5 minutes is how you get banned.
2. HTTP Request node. Set the method to GET, paste your URL, and in the options set:
- Response Format: String (so you get raw HTML, not parsed JSON)
- Headers:
User-Agent: ShedDadTechBot/1.0 (+https://sheddad.tech/bot) - Retry On Fail: enabled, with 3 tries and a 5000ms wait between them
3. A Wait node. Add 2000–5000ms. If you're looping over multiple URLs, this is non-negotiable.
4. HTML Extract node. This is where most people give up. You give it a CSS selector like .product-price and it returns the text. Use your browser's DevTools (right-click → Inspect) to find the selector. Prefer stable attributes like data-testid over auto-generated class names like css-1x9k2p.
5. A Code node for cleaning. Prices come back as £1,299.00 or £1,299.00 inc VAT. Strip the currency symbol and commas with something like:
const raw = $input.first().json.price;
const cleaned = parseFloat(raw.replace(/[^0-9.]/g, ''));
return [{ json: { price: cleaned } }];6. Write to your destination. Google Sheets, Airtable, Postgres, Notion — n8n has nodes for all of them. Add a timestamp column so you can chart trends later.
Handling JavaScript-rendered sites
If the HTML you get back is mostly empty divs, the site is rendering client-side with React, Vue, or similar. Two options:
Find the underlying API. Open DevTools → Network tab → filter by Fetch/XHR → reload. Nine times out of ten there's a JSON endpoint feeding the page. Hit that instead. It's faster, cleaner, and far less likely to break. This is the technique I'd recommend first, always.
Use a headless browser. n8n has community nodes for Puppeteer and Playwright, or you can call a service like Browserless. This is heavier and slower, so only reach for it when the API route is closed off.
If you're running n8n on a Raspberry Pi or a mini PC in the shed, a headless browser will eat your RAM. A used Raspberry Pi 5 with 8GB handles the HTTP-and-parse pattern comfortably; for Playwright you'll want something beefier.
Making it not break next Tuesday
Scrapers break. Selectors change, sites add bot protection, layouts get redesigned. Three habits keep the maintenance burden low:
- Fail loudly. Add an IF node after your extract step: if the price field is empty or null, route to an error branch that emails or Slacks you. Silent failures are the worst kind.
- Log raw HTML on failure. Store the last response body somewhere. When a selector breaks, you'll want to see what the page actually looked like.
- Scrape the least you can. One request that grabs a JSON payload beats five requests parsing five pages.
Also worth knowing: n8n's Error Trigger workflow is excellent for centralising this. Build one error workflow, attach it to every scraper, and you'll get a single notification stream instead of hunting through executions.
If you'd rather skip the build-and-debug phase entirely, I put together n8n Starter Workflows — plug-and-play n8n workflow templates you can import in minutes, including a scraper skeleton with retries, rate limiting, and error handling already wired up. From £9, which is less than an hour of your time is worth.
Hosting and scheduling considerations
n8n Cloud is the easy route — no server admin, automatic updates, and it just runs. Self-hosting on a small VPS is cheaper but you own the maintenance. For scraping specifically, self-hosting has one real advantage: you control the outbound IP, and you can rotate it if needed (though rotating IPs to evade blocks is ethically murky and I'd avoid it).
One UK-specific note: if you're scraping sites that serve different content by region, make sure your n8n instance's egress IP is where you expect. A VPS in Frankfurt will see German pricing, not UK pricing.
For storage, Google Sheets is fine for a few thousand rows. Beyond that, push to Postgres or Airtable. Sheets gets slow and the API rate limits will bite you around the 10,000-row mark.
Wrapping up
Scraping with n8n is genuinely one of the highest-leverage skills you can build as a side-hustler. A scraper that tracks competitor prices, monitors a niche job board, or aggregates local property listings can be the seed of a real product. Do it politely, do it legally, build in retries and error alerts, and prefer JSON endpoints over HTML parsing wherever you can. Start small, run it on a schedule, and let it quietly accumulate data while you get on with your day.
SEO Content Briefs with ChatGPT: A Step-by-Step Workflow
Stop staring at a blank doc. Here's a practical, repeatable workflow for building SEO content briefs with ChatGPT that writers actually want to use.
AI Prompts for Newsletters People Actually Open
Most newsletters die in the inbox graveyard because the subject line is dull and the first line is worse. Here are the exact AI prompts that fix both — plus the rest of the email.
How to Repurpose One Blog Post into 10 Pieces of Content
Write once, publish everywhere. A practical UK-focused workflow for turning a single blog post into ten distinct pieces of content, with prompts you can copy and paste.