Scrapy jobs
I’m running several CPA campaigns and I need a fresh batch of email contacts that are guaranteed to pass Debounce (or an equivalent) with a very low bounce rate. The focus is strictly on email leads sourced from two places: reputable business directories and we...for me: • You pull the data ethically and accurately, respecting each site’s terms. • Every address is verified through Debounce (or a comparable validator) before delivery. • Final output arrives in a clean CSV or Excel file with, at minimum, email, source URL, and any publicly available name or company field you can capture. If you already have an efficient scraping workflow—Python + Scrapy/Selenium, Octoparse, Apify, etc.—and can show sample rows that hit 0% invalid on Debo...
...can feed it straight into my analysis pipeline. I’ll share the exact fields during kickoff, but the scraper must be flexible enough to handle common article elements—headline, body text, author byline, publication date, and source URL—and easy to extend if I add more outlets later. Time is critical. Delivery within 24–48 hours is preferred, so please lean on a proven stack such as Python with Scrapy/BeautifulSoup, Node with Cheerio, or any robust alternative you already master. The script should: • Rotate user agents and accept a proxy list to avoid blocks • Log failed requests for easy reruns • Be clearly commented and organized so I can update selectors myself Deliverables 1. Executable script or notebook with all dependencies not...
I already have a Python-based Scrapy project and now need dedicated spiders that can harvest data from the mobile apps Rabbitmart (Egypt), Voo (Egypt) and Oscar (Egypt). The crawlers should pull every publicly available endpoint (or reverse-engineered one) required to deliver: product details, user reviews, pricing information, the underlying product IDs, any active promotions or discounts and—where the apps expose it—current stock levels. All captured information must be written to clean, well-structured JSON files so that I can feed it straight into my existing pipeline. I’m still unsure whether the apps will demand login or token-based authentication; therefore, please build the spider with the flexibility to plug in credentials, headers or session cookies if ...
Website Content Scraping Required I need content to be extracted from a website and organized in a structured format. The task includes scraping text, images (if required), and other relevant information while maintaining accuracy and proper formatting. Requirements: - Extract...accuracy and proper formatting. Requirements: - Extract content from the specified website. - Preserve headings, paragraphs, and content structure. - Organize the extracted data in Excel, CSV, or Word (as required). - Ensure the data is clean, complete, and free from duplicates. - Deliver the project within the agreed timeline. Experience with web scraping tools (such as Python, BeautifulSoup, Scrapy, Selenium, or similar) is preferred. Please mention your approach, estimated timeline, and cost in your...
...Announcement, Specialty, Experience, Experience, Book on, summary, badges and designations, (Clinic schedule, Fee, Clinic Name, Clinic Location,) of all clinics they have. Education, Med School, Residency, Fellowship Training, Certifications, online clinic hours and availability, Online clinic fee, Affiliations You may harvest the information with the tooling of your choice—Python (BeautifulSoup, Scrapy, Selenium), R, or another reliable stack—so long as the final file imports seamlessly. I am flexible on the exact format (Excel, CSV, or database dump); let me know what works best for your workflow and I’ll confirm before we start. Accuracy is critical. I will verify that every profile on the site is represented and that each required field is populat...
We're looking for an experienced Python developer to stabilize and enhance a production web scraper that's experiencing Cloudflare blocks, broken session handling, and WebSocket instability. The goal is to implement a robust, self-healing scraping pipeline with proper Redis integration. Requirements: - Strong Python experience with web scraping frameworks (Playwright, curl-cffi, Scrapy, or similar) - Hands-on experience bypassing Cloudflare challenges using TLS fingerprint matching and stealth techniques - Redis experience including session storage, TTL management, and pub-sub patterns - Experience with job queuing libraries such as BullMQ or RQ - WebSocket client implementation including reconnection logic, heartbeat management, and binary/JSON frame parsing - Ability to...
...web-scraping specialist who can reliably pull data from Facebook and Instagram. The focus is on public, real-time content; I want clean, structured results that can be analysed right away without additional tidying on my side. To be sure your approach works, please show me a short sample first—10-20 recent records from each platform will do. I am happy with any proven stack (Python + BeautifulSoup/Scrapy, Node + Puppeteer, Selenium, API work-arounds, etc.) as long as you can demonstrate stability, speed, and respect for each platform’s rate limits. Deliverables (all items required): • One working script or tool that scrapes Facebook and Instagram as agreed • A sample dataset for my review before full engagement • Clear instructions or a brief READ...
...would like the final dataset delivered as a CSV. The CSV must include every visible product together with the key details you typically find on a catalogue page—name, price, image URL, description, SKU or ID, and category path. I do not need user reviews, contact details, or any ongoing scheduling; this is strictly a single execution job. Your script can be in Python (requests + BeautifulSoup, Scrapy, or Selenium if the site is JavaScript-heavy) or another language you are comfortable with, as long as it runs reliably on a standard desktop environment. Deliverables: • The CSV file with all product rows and clearly labeled columns. • The executable script or notebook, well-commented so I can rerun it later if the site structure remains unchanged. Pleas...
I want to build a full-featured price comparison tool similar to the core engine behind BuyHatke. The focus is strictly on comparing prices in real time, not on coupons or browser extensions. The system must work seamlessly on the web and be delivered as native iOS and A...pipeline with error handling • Relational or NoSQL store optimised for high-volume price snapshots • Clear documentation and a brief hand-off session I’ll test by verifying that identical products pulled from at least ten merchants show accurate, up-to-date prices across all three platforms within the same refresh cycle. If you’ve tackled price aggregation before—especially with tools like Python-Scrapy, Node, Firebase, or similar—you’ll be able to move quickly. Let&rsq...
...for a detail-oriented web-scraping specialist who can pull data from virtually any public-facing website and hand it back to me neatly organised in an Excel workbook. The site type and data fields will vary from project to project—sometimes it might be product listings, other times articles, reviews, or contact details—so adaptability and solid experience with tools such as Python (BeautifulSoup, Scrapy, Selenium) or equivalent are essential. What matters most is accuracy, clean formatting, and a repeatable process I can rerun in the future. Along with the finished .xlsx file, please include either the script or clear documentation of your method so I can update the scrape if the source site changes. If you have questions about pagination, login barriers, or large d...
I'm seeking a skilled Python developer to work on a project. The specific type of project, primary function, and preferred libraries/frameworks are currently undecided. Ideal Skills and Experience: - Proficiency in Python - Experience with web applications, data analysis, or automation - Familiarity with Django/Flask, Pandas/Numpy, or Scrapy/BeautifulSoup - Strong problem-solving skills - Ability to work independently and meet deadlines Please provide relevant experience and a brief project approach in your bids.
I’m kicking off a Python-based project that will either evolve into a lightweight FastAPI service or a robust web-scraping pipeline—whichever proves the better fit once we start prototyping. Clean, asynchronous code, thoughtful error handling and respect for rate limits are non-negotiable, whether you lean on FastAPI’s dependency-injection patterns or a scraping stack such as Scrapy, Playwright, Selenium, or BeautifulSoup. Before we dive into technical details, I’d like to see concrete proof of your expertise. Please share past work: a GitHub repo, a running demo, or any code snippets that highlight your mastery of FastAPI endpoints, background tasks, pydantic models, or large-scale data extraction from dynamic sites. Real examples will help me understand yo...
...across the sites. Please deliver the data in a single Google Sheets workbook, with a dedicated tab for each source site and clear column headers. So I can plan my budget, let me know your price per site along with an estimated turnaround time once you see the list and any anti-scraping measures that might require work-arounds. • If you already have tooling or scripts in Python, BeautifulSoup, Scrapy, Selenium, or similar, feel free to mention it—speed and reliability matter more to me than the specific stack....
...from a company’s website when that adds value. Core needs • A script, API, or lightweight app that inputs a keyword or industry and returns clean profile data (name, role, company, public URL). • Smart filtering to remove duplicates and obvious non-prospects. • Export options—CSV or JSON at minimum—for easy hand-off to marketing systems. Tech is up to you; Python (BeautifulSoup, Selenium, Scrapy) or a comparable stack is welcome as long as it respects LinkedIn’s limits and complies with website terms. Please send a gedetailleerd projectvoorstel outlining: 1. Your approach to bypassing anti-scraping measures without violating TOS. 2. Key milestones from prototype to final delivery. 3. Examples of similar scraping or data-collection...
I have a stream of numerical information being pulled automatically from several news sites, and I need that raw output cleaned, coded, and d...parses each scrape, identifies the key figures buried in the articles, applies consistent codes to them, and drops everything into a structured file (CSV or JSON works for me). You’ll receive: • the current scraping script and a sample of the raw dump • a field dictionary showing how each number should be labeled or categorised I’m expecting your returned script (Python preferred—BeautifulSoup/Scrapy plus pandas is perfect, but use what you like) along with the final processed dataset and a brief read-me so I can rerun or extend the pipeline later. Accuracy of the coding and reproducibility of results will be ...
...reliable, detail-oriented, and capable of handling various Python-related tasks independently. Requirements: Strong knowledge of Python Experience with Web Scraping and Data Extraction API Integration and Automation Data Processing and Data Management Ability to troubleshoot and optimize existing scripts Good communication skills Ability to meet deadlines Preferred Skills: Selenium BeautifulSoup Scrapy Pandas Requests Flask or Django (optional) Database experience (MySQL, PostgreSQL, SQLite) What I Offer: Long-term work opportunities Multiple projects every month Clear requirements and communication Prompt payment for completed work Please include: Your Python experience Examples of previous projects Your hourly rate or fixed-price expectations Your availability I am loo...
...straightforward: crawl each page, parse the HTML for the specific elements I’ll identify (headings, paragraphs, and a couple of custom tags), normalise any odd characters, then bulk-insert the results so the database is immediately query-able. A repeatable solution matters because I’ll be running the same process weekly as the sites update. I’m comfortable if you build the scraper in Python—BeautifulSoup, Scrapy, or Selenium are all fine—or you can propose another language or library you prefer, as long as it reliably handles pagination and throttles requests to stay respectful of the hosts. Deliverables: • A well-commented script or small codebase that performs the crawl, parse, and SQL insert in one run. • A SQL file (or direct push...
I need a cle...need a clean, one-time scrape of 5,000 products from 1688.com. The final file must be a CSV that contains three reliable fields for every item: the live product link, the exact title as it appears on the site, and the current price. Because 1688 is Chinese-language and often requires dynamic loading or logged-in sessions, please use whatever stack you are most comfortable with—Python, Selenium, Scrapy, Playwright, or another proven crawler—to ensure every row is complete and no data is blocked or throttled. Deliverable • A single CSV (UTF-8) containing 5,000 rows with columns: Title, Price, URL. I will verify by spot-checking random entries against the site, so accuracy and duplicate-free results are essential. Once the file passes that check the...
My day rarely looks the same twice, so I need a flexible side-kick who can switch effortlessly between deep-dive research, fast data scraping, AI-powered task execution, and polished document preparation. One hour you might be pulling product data with Python, BeautifulSoup, or Scrapy; the next you could be steering ChatGPT or another LLM to summarise findings, draft a proposal, or even spin up instructions for a third-party service I delegate to. Core responsibilities • Research – everything from quick market snapshots to more academic or product-level analysis, depending on what the week demands. • Data scraping & cleanup – locating reliable sources, extracting the essentials, and presenting them in clear, reusable formats (CSV, Google Sheets, Airtab...
...exactly as displayed on eBay • current Buy-It-Now or winning price in USD, numeric only The focus is broad—any taxidermy category or sub-category qualifies, so long as the price threshold is met. Please skip any item that cannot supply the minimum three images or any of the other required fields; I prefer completeness over sheer volume. Process details are entirely up to you (Python + Selenium/Scrapy, Octoparse, etc.) as long as the final spreadsheet opens without macros and includes clean hyperlinks and text with no hidden HTML tags. Pay particular attention to pagination and “completed listings” filters to avoid duplicates and out-of-stock items. Before I award the project, share a small sample: five to ten rows in the exact column order above so I c...
I need a reliable way to collect genuine U-S-A mobile contacts—each record must include the person’s name, cell numb...with each site’s TOS. Deliverables • A working web application with secure login and dashboard • Scraper modules for the three platform categories above, with easy ability to add more later • Structured output: name, phone, street address, city, state, ZIP • Documentation covering setup, usage, and adding new sources I’m open to the tech stack you feel most comfortable with—Python (Scrapy/BeautifulSoup), Node.js, or similar—as long as it can handle rotating proxies, CAPTCHA solving, and basic anti-ban tactics. Please outline your proposed approach, past experience with large-scale scraping, and an...
...addresses, SSNs, medical record numbers, account details, etc.). • Focus domains: Healthcare, Finance and Education. • Geographic emphasis: North-American publications at this stage (government portals, open-data sites, public court documents, regulatory disclosures, etc.); other regions may follow once this tranche is complete. • Methods: a mix of automated scraping (Python, BeautifulSoup/Scrapy/Selenium or similar) and classic desk research to reach sources that resist automation. • Output: an organised folder structure plus a spreadsheet/JSON catalog listing document title, source URL, date accessed, domain tag, and a short note of the specific PII fields present. Acceptance criteria 1. Minimum 250 unique documents, balanced across the three doma...
I need a complete, accurate scrape of every dealership listed at https://www.autoscout2... • Contact person’s name • Contact person’s email The finished dataset should be delivered as a clean, de-duplicated CSV file ready for import. Quality is essential: no missing rows, no placeholder values, and emails must be verified as coming from the official source. I will spot-check several records before issuing final approval. If you intend to use Python (BeautifulSoup, Selenium, Scrapy, etc.) or another toolset, that’s fine—just be sure the scraper respects pagination, German characters, and rate limits so nothing is skipped or blocked. The price is fixed; please confirm you can supply the complete CSV with all 19,843 dealerships and the extra...
...record. Here’s how I picture the workflow: • You run an automated scrape on the two sites, bypassing basic anti-bot measures without overloading their servers. • You clean and deduplicate the results so every row represents a unique listing. • You deliver a single, well-structured CSV file (UTF-8) containing three columns: Name, Address, Phone. I’m open to the toolset you prefer—Python, Scrapy, Selenium, or an equivalent stack—as long as the output remains consistent month after month. If you have an existing script that can be adapted, great; otherwise please factor initial script creation and monthly execution into your offer. Reliability is key: the scrape should grab everything visible on the target pages, capture any new records...
I need contact details scraped from 1-2 matrimony sites for my matrimony venture. Requirements: - Experience in web scraping - Familiarity with matrimony sites - Ability to deliver data in organized format Ideal Skills: - Proficient in Python or similar languages - Knowledge of scraping tools like BeautifulSoup or Scrapy - Attention to detail and data accuracy Please provide samples of previous work and estimated delivery time.
I need contact details scraped from 1-2 matrimony sites for my matrimony venture. Requirements: - Experience in web scraping - Familiarity with matrimony sites - Ability to deliver data in organized format Ideal Skills: - Proficient in Python or similar languages - Knowledge of scraping tools like BeautifulSoup or Scrapy - Attention to detail and data accuracy Please provide samples of previous work and estimated delivery time.
I need all...description, availability and any visible category tags—into the right columns. Everything is public and loads in plain HTML, so no login is required, but I do want the pull to be 100 % complete, free of duplicates and ready for quick analysis the moment I open the file. Deliverables • One .xlsx file containing the full product list • The reusable scraping script (Python with BeautifulSoup, Scrapy, or a similar tool) or a clear step-by-step method so I can repeat the extraction later Let me know your estimated turnaround and any potential hurdles you see with pagination or rate limits so we can address them up front. In total it there are 160.000 product entries with 5 specifications that I need in an excel or csv file. So its a quick job ...
...team can sort and filter with ease. At minimum, every row must include a working email address; if you happen to capture basic contact details (company name, phone, street address) while you work, feel free to include those in adjacent columns, but the email is the critical field. Before delivery, validate each address to avoid hard bounces or obvious spam traps. Any toolset is fine—Python, Scrapy, Hunter, NeverBounce, or similar—so long as the final file opens smoothly in Excel and every address you provide passes a quick test send. I’ll review by spot-checking twenty random entries: if they bounce or reach role-based accounts that commonly trigger filters (info@, sales@, support@), I’ll ask for replacements. Let me know how quickly you can complete...
...that pulls specific text content from a particular website each day and appends the results to a clean, well-structured CSV. The crawl has to run automatically on a 24-hour schedule, capture only the text I specify, and overwrite or update the file so I can download it at any time. Please handle the full workflow—from parsing the HTML through a stable method (Python with Requests + BeautifulSoup, Scrapy, or a comparable stack) to setting up the daily trigger (cron job, cloud function, or Windows Task Scheduler). The script must cope gracefully with minor layout changes and alert me if anything breaks. Deliverables • Fully commented source code • One sample CSV generated from a test run • Setup notes so I can deploy the job on my own server Acceptan...
...clean, the same system should launch automated outreach campaigns—email sequences, LinkedIn touches, or WhatsApp nudges—while an AI tracker watches for local construction triggers such as new CREDAI member additions or early-stage excavation alerts. Every interaction and trigger should feed back into my CRM so I can see hot prospects in real time. Deliverables • A deployed agent (Python/Node, Scrapy/Selenium, LangChain or similar) that scrapes the three source categories safely and respects rate limits. • Contact-enrichment pipeline with validation ≥ 95 % accuracy. • Automated outreach workflows pre-loaded in my CRM/marketing-automation tool, ready to run. • Trigger dashboard or webhook that flags new projects within an hour of detection....
...minimum I expect company name, full postal address, phone, email, and any registration numbers available on the SOS records. • Data hygiene – deduplicate across sources, normalise addresses, and flag obviously invalid phones or emails. • Delivery – a single, well-structured CSV file ready for immediate import. Alongside the file, include the scripts or notebooks you used (Python + BeautifulSoup/Scrapy/Selenium or similar) and a short README so I can rerun the pipeline later. Acceptance criteria 1. At least 95 % of rows must contain all mandatory fields. 2. No more than 2 % duplicate businesses when matching on name + address. 3. Scripts re-create the same CSV on a fresh machine with only the listed dependencies. If this matches your expertise i...
...category/type Daily/weekly/monthly pricing Rental rates and fees Taxes or additional charges (if visible) Availability information Pickup and return locations Promotions or discounts Vehicle specifications/features Images/URLs (if applicable) Any other relevant offer details displayed on the website Technical Requirements Preferred technologies: Python Selenium and/or Playwright BeautifulSoup, Scrapy, or similar libraries are acceptable if needed The scraper should: Navigate dynamically loaded pages Handle pagination and multiple locations Work reliably across all available branches/locations nationwide Be modular, clean, and maintainable Include error handling and logging Avoid duplicate records Generate structured output files automatically Deliverables The final delivery must ...
I need 3,000 fresh business-directory records pulled into a clean spreadsheet. Each row must include: • Company name • Full address (street, city, state, ZIP) • Primary phone number • Contact email The source is a public online directory; I’ll provide the exact URL and any search parameters as soon as we start. Please use an automated method (Python, BeautifulSoup, Scrapy, Selenium … whatever you prefer) that respects rate limits and captures the data accurately—no missing fields or obvious duplicates. Deliver the final file in CSV or Excel and include the script you used so I can rerun it later if needed. I’ll review random samples against the live site; records must match 98 %+ for acceptance. Let me know your estimated tu...
I need a clean, well-commented Python script that starts from a single root URL, follows every internal link it finds (no keyword or structural filtering at all), and extracts only the visible text content on each page. The script should rely on mainstream libraries—requests plus BeautifulSoup is fine, but feel free to propose Scrapy or an async stack if it fits better. Please keep the code modular so I can later drop individual functions into a bigger application. Core expectations • Crawl every reachable link within the domain, respecting and an adjustable polite delay. • Skip images, PDFs, or other binary assets; focus strictly on textual information. • Save each page’s URL alongside the extracted text in a single output file (CSV or JSON&mda...
...directories, and recognised industry publications. Please harvest the data accurately, avoid duplicates, and verify that every record is current. Deliverables • One Excel workbook with separate, clearly labeled sheets (or columns) for company details, product lines, and contact information. • A short note explaining any assumptions, filters, or automated tools/scripts (Python, BeautifulSoup, Selenium, Scrapy, etc.) you used so I can reproduce or update the crawl later. • A quick quality check summary highlighting any gaps you could not fill and why. I’d like the first sample of 25 companies within a few days so we can confirm structure before you proceed to the full scrape. Let me know your estimated turnaround time and any questions you have about data...
We have an Excel file containing 4,300+ company websites. You need to populate only 1 Column head namely "Teas Details". I need a Python-based automation solution (Playwright, Selenium, Scrapy, etc.) that will: 1. Visit each company website. 2. Look for Team, Leadership, About Us, Meet the Team, Our People, Management, Staff, or similar pages. 3. If a team/leadership page exists, count the visible team members listed on that page and populate the "Team Details" column with the numeric count only (examples: 2, 5, 12, etc.). 4. If no team page is found, populate: No Team Page Found 5. If the website field is blank or missing, populate: No Team Page Found Requirements: * Process 4,300+ records automatically. * Update only the Team Details column. * Return...
I run a regional event directory that follows more than 400 venue URLs. At the moment I rely on Manus AI, which only captures about 60-70 % of what is published. I want a purpose-built scraper that raises that coverage to...Acceptance criteria • Test run across the current 400 URLs shows ≥95 % capture of unique events. • No duplicate entries created during the same run or across consecutive runs. • All fields (title, date/time, venue name, source URL) populate correctly in my database. • Runtime per full scan stays under two hours on a 2-vCPU VPS. If you have deep experience with Python scraping frameworks (Scrapy, Selenium, Playwright), PDF parsing libraries, and API integration, I’d love to see an outline of how you’d tackle the mix of f...
I need a script/app that can pull seller and product information from an online marketplaces. The scraper should be able to capture standard item details, handle pagination, and direct seller-URL or seller-ID input. Selenium, Scrapy, Playwright, or a well-documented REST API wrapper are all acceptable approaches. Keep rate-limiting, captcha handling. If you have relevant examples or a demo snippet, feel free to share them in your bid.
...one shot. Key points you should know: • The site is public but rate-limited, so the script must handle pagination, delays or rotating headers/IPs if necessary. • Output should be a clean CSV or Excel file containing at minimum two columns: Email and Phone. If the same record appears multiple times, the script must de-duplicate it automatically. • I would like the finished Python code (Scrapy, BeautifulSoup, Selenium, or any other proven library) so I can rerun it whenever the database grows. Please include a brief README explaining setup and usage. Acceptance criteria 1. A single data file containing all available email and phone entries from the site. 2. Script runs end-to-end on my machine without manual intervention beyond providing login or capt...
I'm looking for a skilled Python programmer to develop an automation script for...Requirements: - Develop a Python script to scrape data from the specified tire manufacturer website. - The script should be efficient, reliable, and able to handle potential changes in the website structure. - Data to be collected will include product fitment. The data should be saved in a CSV or XLS format Ideal Skills: - Proficiency in Python and web scraping libraries (e.g., Beautiful Soup, Scrapy, Selenium). - Experience with handling and processing large datasets. - Knowledge of anti-scraping techniques and how to navigate them. - Familiarity with the tire industry is a plus but not necessary. Please provide examples of previous web scraping projects and ensure you can deliver within the sp...
...vary from business directories and company websites to social platforms—and potentially any other page that reveals the information—I’m looking for a developer comfortable mixing and matching techniques: headless browsers for dynamic pages, straight HTTP requests where possible, and smart rate-limiting or proxy rotation to keep everything compliant and undetected. You’ll likely rely on Python (Scrapy, BeautifulSoup, Selenium or Playwright) as well as TypeScript where it makes sense, for example in a Node-based microservice or a lightweight dashboard that lets me trigger new scraping jobs and download the resulting datasets. I don’t mind which framework does what as long as the whole pipeline is clear, well-documented and easy for me to redeploy. De...
...CVs based on years of experience and industry sectors. This is a performance-based gig with payment per valid CV submitted, requiring a minimum output of 15,000 CVs per month. The position is temporary, with an initial duration of 3-6 months, and can be extended based on performance. Key Responsibilities 1. Use web scraping tools and scripts (e.g., Python with libraries like Beautiful Soup, Scrapy, or Selenium) to extract CVs from legitimate public sources such as job portals, professional networking sites (ensuring compliance with terms of service), and open databases. 2. Filter and collect only CVs that demonstrate relevant Gulf region experience (e.g., work history in Gulf countries). 3. Segregate the scraped CVs into categories based on: - Years of experience (e.g., 0-2 y...
...a single CSV file. The scrape must follow GSMArena’s site structure so the file is easy to keep in sync later, and every field that appears on a spec sheet—brand, model, announcement date, display, chipset, memory, cameras, battery, network bands, dimensions, OS and so on—should be included. Please build the script in Python and rely on standard scraping tools such as Requests, BeautifulSoup, Scrapy or Selenium (feel free to combine them if pagination or dynamic content requires it). I want the code and short setup notes alongside the final CSV so I can rerun the process on my own machine when new devices appear. Deliverables • Python source code with clear comments • CSV containing one row per phone and one column per spec field • README with s...
...emails, phone numbers, physical addresses, and active website or profile URLs. • Cross-check each entry across at least two of the approved sources to minimise errors and duplicates. • Organise everything in a clean, well-structured Excel spreadsheet with separate columns for every field, consistent formatting, and no blank placeholders. I’m open to your preferred toolkit—Python, BeautifulSoup, Scrapy, Selenium, or even manual curation if that’s faster for niche sites—so long as the final sheet is accurate and ready to filter or bulk-import. Acceptance criteria 1. Minimum 95 % valid emails (hard-bounce tests welcome). 2. Zero duplicate rows. 3. Excel file passes a quick spot-check against the live sources. If this sounds straightforw...
I need a reliable data-scraping specialist to build a contact list of people in Uruguay who are confirmed credit-card holders. I only require two fields per record—valid full name ...estimated turnaround. • Output: a clean CSV or Google Sheet containing Name and Phone columns (extra fields like occupation are welcome if they come at no added effort). Ethical and legal compliance is non-negotiable—no breaches of platform terms or local privacy regulations. When you respond, include: 1. A brief outline of your scraping approach and preferred tools (e.g., Python, Selenium, Scrapy, Phantombuster). 2. Examples of similar contact or LinkedIn extractions you have completed. 3. An estimated record count and delivery timeframe. I’m ready to start as so...
...Accuracy and completeness of fields such as title, price, images, description, rating, and availability. • Resilience to site changes and anti-bot measures (dynamic content, pagination, CAPTCHAs). • Output delivered in structured CSV or JSON so I can drop it straight into my own pipeline. • Clear instructions on how to run or schedule the scraper on a standard Python environment; feel free to use Scrapy, BeautifulSoup, Selenium, Playwright, or any modern AI/ML technique that helps maintain extraction quality. I’ll provide the exact URLs and the final list of fields after we agree on an approach. If you’ve built scrapers for Amazon, eBay, Walmart, or similar platforms, share a quick example or demo URL so I can see your style. I consider the job done w...
...blocked. Scope of Data Text Data: Product name, complete technical specifications/descriptions, current price, and stock status for each of the ~250,000 products. Images: Primary product images saved locally, ensuring high-quality, watermark-free versions. Metadata: SKU, part compatibility, and source URL. Technical Stack & Requirements Preferred Stack: Python (Playwright or Scrapy with Playwright integration) is preferred for handling dynamic JavaScript content and complex DOM interactions at scale. Output Format: Structured data exported to CSV or JSON (or direct database dump), with images organized in a mirrored folder structure. Documentation: A concise README detailing how to update search parameters and selectors when the site layout cha...
I have a list of online sources where I need all available contact details—names, roles, emails, phone numbers, and any publicly listed social links—pulled into a single, well-structured Excel file. Many of these sites use pagination, dynamic loading, or the occasional bot check, so the solution has to be resilient (Python + BeautifulSoup, Scrapy, or Selenium generally work best in my experience, but I’m open to whatever stack you prefer). Here’s what I expect: • A repeatable script or crawler that targets the specified URLs and captures every relevant field. • Cleaning and de-duplication before delivery so the spreadsheet is ready for immediate use. • An Excel workbook with separate tabs if multiple sites need to stay segmented, otherwise...
I’m building a fresh prospect list and need someone who can automatically pull contact information from a set of business websites I’ll supply. The goal is simple: turn each URL into clean, usable leads that include company name, contact person (if available), email, phone, and any publicly listed address. You’re free to employ Python with BeautifulSoup, Scrapy, Selenium or a comparable stack—the only requirement is that the solution runs reliably, respects reasonable scraping etiquette, and can be rerun whenever I add new domains. I’ll provide the starting list of sites plus a short test batch so we can verify everything is being captured correctly. Deliverables • A working script (with clear setup notes) or a lightweight app that runs on ...
...sub-category path, all image URLs, plus every variant attribute such as colour, size or other options the site shows. The scraper should run on demand and be easy to schedule for weekly updates. Anti-bot countermeasures—rotating proxies, polite timing, and basic CAPTCHA handling—will be essential because some of these sites throttle traffic quickly. I’m comfortable with a Python stack, so tools such as Scrapy, BeautifulSoup, Selenium or Playwright are welcome as long as the final code is clean, well-commented and handed over in a private Git repository. Deliverables • • CSV or Excel export containing: product URL, title, description, category, sub-category, price, variant attributes, image links • Compressed folder (or cloud bucket) of downl...