Subject Matter Expert - Web Scraping & Crawling (Web Data Platform)
Full-time
Hyderabad, Telangana, India
Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.
Role SummaryOwn and extend a production-bound POC that combines resilient web-crawling engineering with an AWS Bedrock-based LLM verification agent, in a healthcare-data-compliance-sensitive context.
Responsibilities- Replace scraped SERPs with a paid search API behind the existing search_domain() interface.
- Move fetch.py to async httpx with per-domain concurrency limits and politeness budgets.
- Introduce proxy rotation and egress management; retire the single-IP failure mode.
- Make failure loud: typed outcomes, structured logs, metrics, drift and volume alerts.
- Reverse-engineer payer XHR/JSON endpoints to replace Playwright recipes wherever possible.
- Honor robots.txt Disallow and Crawl-delay; build a per-domain ToS and licensing register.
- Containerize and schedule the pipeline; add CI running the offline tests on every change.
- Add HAR replay and golden-file parser tests on real payer HTML, plus daily canary crawls.
Web & protocol fundamentals — HTTP/1.1, HTTP/2, HTTP/3, TLS, JA3/JA4 fingerprinting, cookies/sessions, OAuth2/JWT/SSO, DOM, encodings [Must]
Legacy stack (real mileage) — urllib, requests, mechanize, cURL/wget, BeautifulSoup, lxml, html5lib, XPath/XSLT, Scrapy (middlewares, pipelines, CrawlSpider), Selenium 3 patterns, Splash/PhantomJS, RSS/Atom/SOAP/XML feeds, ASP.NET __VIEWSTATE postbacks, frame & table-layout scraping, sitemap.xml, WARC/Common Crawl [Must]
Modern stack — Playwright (contexts, tracing, network interception), Puppeteer, Selenium 4/CDP, httpx/aiohttp + asyncio, curl_cffi TLS impersonation, selectolax, scrapy-playwright, Crawlee/Apify, JSON-LD/microdata harvesting [Must]
Reverse engineering — Private/undocumented JSON APIs, XHR & fetch interception, GraphQL, mobile API capture (mitmproxy/Charles), bundled-JS deobfuscation, WASM challenge analysis [Must]
Methodology breadth — API-first vs browser rendering, static vs JS-rendered SPA (__NEXT_DATA__, Nuxt payloads), BFS/DFS frontier management, URL canonicalization & dedup (bloom filters/hashing), full-refresh vs incremental/CDC crawling, distributed queue-based crawling (Redis/Kafka/SQS), authenticated sessions, pagination & infinite scroll, per-domain politeness [Must]
Non-HTML extraction — PDF (pdfplumber, PyMuPDF — in use today), OCR (Tesseract/cloud), Office formats, LLM/vision-assisted extraction and its cost & failure modes [Must]
Anti-bot & reliability — Cloudflare, Akamai, DataDome, PerimeterX, Imperva, Kasada; browser/canvas/WebGL fingerprinting; residential/datacenter/mobile proxy rotation; CAPTCHA landscape; honeypot detection; backoff with jitter, circuit breakers, idempotent retries [Must]
Data engineering — Python expert (async, typing, profiling), SQL, pydantic validation (in use); pandera, entity resolution, Postgres/Mongo/S3/Parquet, Airflow/Prefect/Dagster, Docker/K8s, CI/CD to introduce; Node/TypeScript useful [Must]
Testing & observability — vcrpy/responses, HAR replay, golden-file parser tests, canary crawls, selector-drift alerting, volume-anomaly detection, structured logging & metrics [Must]
Build vs. buy — Evaluating Zyte, Bright Data, Oxylabs, Apify, ScrapingBee, Firecrawl on cost, coverage, and risk [Preferred]
Legal & ethical — robots.txt and rate-limit adherence, ToS/CFAA awareness, GDPR & PII minimization, licensed-API-first sourcing; designs defensible, low-risk acquisition strategies in partnership with Legal [Must]
Experience — Required- 8+ years in data acquisition; 5+ years owning a scraping platform end to end.
- Proven scale: 1,000+ distinct domains or 10M+ pages/month sustained in production.
- Rescued a brittle legacy scraper, with before/after reliability and cost numbers.
- Mentored engineers; set crawl standards, review practice, and on-call runbooks.
- Degree optional — equivalent practical experience is fully accepted.
- US payer policy, formulary, or prior-authorization document domain knowledge.
- Open-source contributions to Scrapy, Playwright, Crawlee, or a parsing library.
- LLM-assisted extraction at controlled cost — we run AWS Bedrock in verifier/.
- Compliance or legal-review exposure on data acquisition programmes.
Chryselys is proud to be an Equal Employment Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.