Back to all positions

Engineering

Full Stack Web Scraping Engineer

Full-timeRemote

We are searching for a new Badger buddy: a full stack scraping engineer who runs a small army of AI agents (Claude Code, Codex, Cursor, Antigravity) and can tell when one of them is lying.

Wanted: one Badger buddy

ScrapeBadger is looking for a new Badger buddy. Not a "rockstar", not a "ninja". A badger: digs where others give up, is awake when the walls change at 3 a.m., and does not care how big the thing on the other side of the fence is.

Formally, the role is Full Stack Web Scraping Engineer. Informally, you will own scraping targets end to end: from finding the cheapest way in that actually holds, to the API endpoint, the SDKs and the docs that sell it, with a small army of AI agents doing most of the typing.

The non-negotiable: you are an agentic AI power user

This is the most important paragraph on this page, so it gets its own section and a slightly larger font in our hearts.

We do not write most of our code by hand any more. Agents do. The human designs the experiment, reads the production data, steers the agents, and catches them when they lie. Our engineering knowledge base is maintained by an LLM. Every task starts with an agent reading it and ends with an agent writing back what it learned. Skills, hooks, subagents and MCP servers are the tools of the trade here, not a party trick.

So we need someone who is genuinely, unusually good at this. Not "I tried Copilot once". We mean:

  • Claude Code: CLAUDE.md and project skills, custom slash commands, hooks, subagents and parallel worktrees, MCP servers, plan mode, headless runs in CI. You have opinions about context management and you know why a 200-line skill beats a 2,000-line prompt.
  • Codex: AGENTS.md, sandboxed and cloud runs, kicking off several tasks in parallel and reviewing the PRs they open instead of babysitting one terminal.
  • Cursor: rules files, agent mode, background agents, MCP, @-context done deliberately, checkpoints and when to roll back to one.
  • Antigravity: running multiple agents from one manager, browser-in-the-loop verification, artifacts and task plans you actually read before approving.
  • Across all of them: you write skills and rules that make the next run better, you delegate to subagents to keep your own context small, you have built or shipped an MCP server, and you can tell within ten seconds when an agent is confidently wrong. A green checkmark from an agent means nothing to you until you have seen the numbers.

If that paragraph sounds like cheating, this is not the job for you. If it sounds like a normal Tuesday, keep reading.

What you will actually do

Web scraping at scale is a game against people who are paid to stop you. The page is usually the most expensive way in. The job is finding the cheapest way that holds, keeping it holding while the target changes underneath you, and shipping the result all the way to a paying customer.

  • Keep targets yielding against live anti-bot systems. Design small, honest experiments, read the outcome from production data, and move the lever that actually moves.
  • Care about what a request looks like on the wire. TLS and browser fingerprints, header order, client hints, personas that stay coherent. You know when the client is not the problem.
  • Reverse-engineer private and mobile APIs the way a detective reads a room: internal endpoints, inline JSON, GraphQL, signed requests.
  • Run the machinery: browser farms, proxy pools, pacing and retries, the parts nobody sees until they fall over.
  • Build the product around it: the API endpoint, the dashboard playground, the public landing page, the docs, and the SDKs that make a customer's first request work.

How we work

  • The cheapest way in wins, and it is usually not the page. Browsers are the last resort, not the first.
  • Validate the payload, never the status code. A small body with a happy status is the thing to distrust.
  • Change one thing at a time and measure it. Every false breakthrough we ever had came from skipping this.
  • Root cause, not symptom. One ticket usually points at ten endpoints.
  • Production data before code. The numbers answer more than the stack trace.
  • Write it down. Every finding and every dead end goes in the knowledge base, cited and dated. The agents read it before you do.
  • Agents are the toolchain. You direct them fluently and distrust them professionally.

Stack, in one breath

Python and FastAPI on the scraping side, TypeScript and Next.js on the product side, PostgreSQL and Redis underneath, Docker on our own servers, real browsers and low-level HTTP clients for the hard targets, and an agent toolchain of Claude Code, Codex, Cursor and Antigravity with a shared knowledge base and MCP.

Who does well here

Need to have

  • Agentic AI mastery. See the section above. This one is not negotiable and it is checked first.
  • You have shipped and operated scrapers against real anti-bot vendors and can tell us which ones beat you and why.
  • Fluent in async Python; comfortable in TypeScript and React. Your agents write it, you review it like it is yours, because it is.
  • You can read HTTP at the wire and find the real API behind a page with nothing but DevTools.
  • You debug from production data and bring numbers to the argument.
  • Linux and Docker without ceremony.
  • You write clearly. The knowledge base is our institutional memory, and the agents read it back to us.

Nice to have

  • Go or Rust.
  • JavaScript deobfuscation and VM-based anti-bot obfuscation.
  • Mobile reverse engineering (Frida, JADX, Ghidra).
  • You have published a skill, a rules file, or an MCP server that other people use.

Probably not for you if

  • You still type every line yourself and are proud of it.
  • An agent's green checkmark is proof to you. Here the fixture is the first suspect.
  • You want a fully written spec. A ticket here is a customer's sentence and a page in the knowledge base. The spec is what you (and your agents) find.
  • You prefer a lane. The same week can hold a reverse-engineering session, a billing webhook and a landing page.

How hiring works

  1. Send a note. One scraping problem you solved, what you measured, and how you drove your agents through it. A transcript or a skill you wrote beats a CV.
  2. A 60-minute call. We walk through one of our real incidents together, with the logs. Bring your agent of choice.
  3. A paid, time-boxed take-home. Half a day against a live target, done with your agents. We look at the result and at how you steered. Paid whether or not we continue.
  4. Offer. Usually within two weeks of the first note.

Details

  • Location: Remote. We overlap on European hours.
  • Engagement: Full-time, or a long-term contract.
  • Start: As soon as you can.
  • Legal and ethics: We collect publicly available data, and we care where our proxies come from.
  • Equipment: Your own machine, and yes, we pay for the agent subscriptions.

Apply below. Bring your agents. We will handle the CAPTCHA, which, given what we do all day, is the least we can do.

Apply

One problem you solved, what you measured, and a link to something you built. That is the whole application.