A guy who sells phone cases on Amazon called me last year convinced someone inside his company was leaking pricing to a competitor. The competitor was matching him within hours of every change, always one dollar less. I told him nobody was leaking anything, his competitor was almost certainly running a price scraper.

He went quiet for a second and then asked me two questions back to back: “is that legal” and “can you do it to them.” I said probably and yes. That phone call is how I got into this work and it’s the story I’m going to use to explain everything in this article, because every concept in price scraping makes more sense when you see it through a real project.

What Is Price Scraping?

The phone case client was spending about two hours every morning checking what his competitors charged. He’d open each listing on Amazon, write the price in a Google Sheet, compare it to his own, decide whether to adjust.

Every morning. By the time he finished all 200 products, some of the prices he’d checked first had already changed again because Amazon pricing moves fast in his category. He was doing this by hand and losing the race every day.

What is price scraping? It’s what replaced that two-hour morning routine. I called my friend Derek on a Saturday and told him what I needed. He said “oh, a scraper” like it was obvious. We wrote a script over the weekend, about 30 lines of Python, that visits each competitor’s product page, finds the price on the page, and saves it.

Instead of my client opening 200 tabs and squinting at numbers, the script does it in about four minutes. The industry calls this price intelligence, which Derek says sounds like a CIA program. He’s not wrong. It’s really just automated competitor price checking.

How Does Price Scraping Work?

Derek came over with his laptop and a six-pack of Topo Chico and I mostly watched him work. He’d ask me for a URL, paste it into the script, run it, and a number would appear in the terminal. “That’s the price,” he’d say. He did this a couple times to show me it worked and then pasted in about 200 URLs from a spreadsheet and ran the whole batch.

Then he spent maybe twenty minutes setting up two other things I’d never heard of, one for saving the data and one for making it run automatically every six hours. PostgreSQL for the storage, I remember that because the name is absurd. The timer one he described as “an alarm clock for code” which stuck with me better than its actual name.

Monday morning I opened the database and it had a whole weekend of competitor prices for all 200 products. The phone case guy sent back seven exclamation marks, no words.

That’s when I knew product price scraping was going to be a real service and not just a weekend project. Here’s roughly what Derek’s script looked like, because people keep asking me for the code and it’s almost embarrassingly short:

    import requests
    from bs4 import BeautifulSoup
    
    url = "https://www.example.com/product-page"
    page = requests.get(url)
    soup = BeautifulSoup(page.text, "html.parser")
    
    price = soup.find("span", class_="price").text
    print(price)  # $17.99

That’s it. That’s the whole concept. Find the page, find the price element, grab the text. Derek had it looping through a list of URLs and saving each price to PostgreSQL with a timestamp, but the core logic is those four lines.

I remember looking at it on his screen that Saturday and thinking I’d wasted a year of my career not knowing this existed. The script ran beautifully for ten days. On day eleven I checked it and it had been frozen on an Amazon CAPTCHA since 3am.

Just sitting there, staring at a grid of traffic light images, waiting for someone to click them. There was nobody to click them. The price scraper doesn’t have hands.

Ecommerce Price Scraping Use Cases

The phone case guy was my first client and his use case is the most common one I see: tracking competitor prices on Amazon to match or undercut them faster. He gets a Slack ping now whenever a competitor drops more than ten percent and adjusts his own pricing during lunch. He told me last month it’s saved him roughly eight grand in sales he would have lost from catching price drops a day late.

The skincare brand was my second client. Sixteen competing Shopify stores, about 640 price points daily. Their marketing director, Laura, wanted to know when competitors changed strategy. The scraper caught a competitor dropping their entire catalog 15% overnight.

Laura called me sounding excited, which for Laura is unusual. She asked if they should match. I said wait. A week later the competitor reversed everything. Without the scraper Laura said they would have panicked into a matching strategy for a price drop that lasted exactly seven days.

A travel agency I picked up after the skincare brand uses the same infrastructure for web scraping for price comparison across three booking platforms. 300 flight routes, checked every twelve hours.

Different product pages, same technical setup. Ecommerce price scraping, flight price scraping, the underlying engineering is identical once you solve the getting-blocked problem. Which brings me to the walls I hit.

Challenges of Price Scraping

I hit four walls in three months and each one almost made me give up. Wall number one: CAPTCHAs. Amazon throws them after about 150 to 200 requests from the same IP. My scraper froze on day eleven and sat on a traffic light grid for six hours before I noticed. I called Derek. He laughed, which was not the response I wanted, and said I needed different IPs. I said I only had one. He said that was the problem.

Wall number two: I bought datacenter proxy IPs. Worked for three days before Walmart blocked the IP entirely. Not CAPTCHAd, blocked. 403 on every request. Bought another. Five days. Third one, four days. Derek asked me how many I was planning to burn through. I didn’t have a good answer.

Wall number three took me the longest. Half the Shopify stores I was scraping for the skincare brand returned blank prices. I spent three days convinced I had a bug. Called Derek on Saturday.

He opened a blank page, clicked View Source, scrolled for thirty seconds, and said “these prices get added by JavaScript after the page opens, you’re photographing a whiteboard before anyone writes on it.” He told me to use Puppeteer. I asked what that was. He said “Chrome with a remote control, give me your laptop” and set it up himself in ten minutes.

Wall number four was the sneakiest. Everything else was working and Amazon kept randomly blocking me once or twice a week. Derek opened a site called BrowserLeaks on my laptop and showed me that all my scraping sessions had the same canvas fingerprint. Different IPs, same browser identity. He said “Amazon can see these are all the same machine” and I said “why didn’t you tell me that two months ago” and he said he thought I knew.

Why Proxies Are Used for Price Scraping

Every wall I hit came from the same root problem: the website figured out too many requests were coming from one source. Derek’s solution for wall one and two was residential proxies from Floppydata.

I asked what the difference was between those and the datacenter ones I’d been burning through, and he went into this whole explanation that I mostly zoned out during, but the gist was that Amazon keeps a list of IP addresses from hosting companies and blocks them on sight, whereas IPs from regular household internet connections look the same as a normal shopper.

I signed up for Floppydata, plugged an IP into my scraper, and the CAPTCHA that had been sitting on screen for three hours just disappeared. The page loaded. I stared at it for a second because I genuinely expected another CAPTCHA behind the first one. There wasn’t.

Plugging a proxy into the script was about two extra lines. Derek showed me this after I asked him to write it down because I knew I’d forget:

    proxies = {
        "http": "http://user:[email protected]:8080",
        "https": "http://user:[email protected]:8080"
    }
    
    page = requests.get(url, proxies=proxies)

Same script as before, just with a proxies argument on the request. The URL goes through Floppydata’s server instead of my apartment wifi and Amazon sees a residential IP from wherever Floppydata routes it instead of my home address in Austin that it had already flagged.

Derek said the important thing was rotating these, not using the same one for every request, so I ended up with a list of about twenty IPs and the script picks one at random for each product page. The CAPTCHAs stopped completely.

For the fingerprint problem, wall four, Derek had been using 1Browser for six months without telling me, which I’m still annoyed about. He created three browser profiles on my laptop, plugged a different Floppydata IP into each one, ran BrowserLeaks again.

Three different hashes from the same machine. He connected them to my Puppeteer scripts through the automation thing and told me to run the Amazon scraper overnight. No blocks. Next morning, no blocks.

A week went by and I texted Derek “I think something’s wrong, the scraper hasn’t failed in a week” and he said “nothing’s wrong, it’s just working” which was a weird sentence to hear after two months of constant failures.

The setup now: Floppydata for residential IPs, about $40/mo across all my projects. 1Browser free plan, ten profiles. One client has a team of three on Gologin for cloud profile sharing, $24/mo, though I make them use Floppydata IPs because Gologin’s shared pool has been touched by every Gologin user on earth.

Is Price Scraping Legal?

Every client asks this eventually. The phone case guy asked during our first call. The skincare brand’s legal department wouldn’t approve the project until they had something in writing. So I called a lawyer friend named Elena who charged me for the consultation, which I thought was rude between friends but apparently is normal if you’re a lawyer.

I said “my client wants to automatically check competitor prices on public websites, is he going to get sued.” She said “probably not.” I said “can you be more specific.” She said there was a Supreme Court case in 2022 and started explaining it and I said “Elena, I need a yes or a no.” She said “it’s a yes with footnotes” and sent me a two-page email that I forwarded unread to the skincare brand’s legal team. They approved the project. The phone case guy just said “cool.”

Three years of ecommerce price scraping and the closest I’ve gotten to trouble is a cease-and-desist email that turned out to be a template some law firm sends to anyone whose scraping shows up in their server logs. I replied. They didn’t. Elena asks me every few months if I’ve been sued yet. I haven’t. She sounds mildly disappointed, which I’ve decided not to take personally.

Getting Started with Price Scraping

A guy named Raj messaged me on LinkedIn last month about price scraping for his supplement brand. I wrote back a long message and then deleted it and wrote a shorter one: “pick one competitor, ten products, write a script, run it from home, don’t buy anything yet, wait for it to break, then message me again.”

He messaged me four days later: “Amazon keeps giving me CAPTCHAs.” I sent him a link to Floppydata and said “wall number one.” He wrote back a week after that asking why some Shopify stores return blank prices. I said “wall number two, get Puppeteer.” I’m waiting for wall number three. He doesn’t know about the fingerprint thing yet. When he asks I’ll send him this article.

Derek texted me this morning while I was writing this. He said “did you give me credit” and I said “you’re in basically every paragraph.” He said “good” and then a minute later sent a follow-up that said “did you mention the thing about how the scraping part is easy” and I said yes Derek I mentioned the thing. He said “good because I told you that on day one and you spent three months ignoring me.” Raj will probably listen faster than I did. Most people do. Derek sends his regards.