
Closed
Posted
Paid on delivery
Build a geocoded USPTO patent dataset (submarine cable technologies) — data collection + cleaning + geocoding Overview I'm building a structured dataset of US patents in a defined technology area (submarine fibre-optic cable systems and related optical/materials technologies) for a quantitative analysis of where this innovation happens and has happened geographically. The output feeds an academic project modelled on recent work in the economics of innovation that maps patenting to geography — see [attached paper / link], e.g. the commuting-zone patent maps in Figure 13. That's the kind of geocoded output I'm after. What you'll build Pull US patent records for a set of search terms and CPC classes I provide, from PatentsView, Google Patents (BigQuery), or USPTO bulk data — whichever you can justify. For each patent: patent number, title, abstract, filing and grant dates, assignee(s), inventor(s), and inventor/assignee locations. Clean and disambiguate assignee names (e.g. collapsing "Corning Inc" / "Corning Incorporated" / "CORNING INC." into one entity). Geocode inventor and assignee locations to city-level lat/long, preserving the raw address components (city, state, country as separate fields) so locations can later be mapped to US commuting zones. Deliver as CSV + Excel, plus the Python or R script so I can re-run and extend it. Workflow Start with a sample of a few hundred patents, fully processed end-to-end, so we confirm the pipeline is correct before scaling to the full corpus. I'd rather get the sample exactly right than receive a large, messy dataset. To apply, tell me briefly: Which patent data source you've used, and on what specific project. How you'd geocode inventor addresses to city level, and how you'd handle missing or malformed location data. How you'd approach assignee name disambiguation. A link to a past data project counts for far more than a written description. Exposure to the geography of innovation or innovation economics is a strong plus but not essential — clean, reproducible data work is what matters most. Timeline The dataset build needs to be completed between now and 25 September 2026.
Project ID: 40667917
102 proposals
Remote project
Active 6 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
102 freelancers are bidding on average £441 GBP for this job

⭐⭐⭐⭐⭐ Build a Geocoded USPTO Patent Dataset for Submarine Cable Technologies ❇️ Hi My Friend, I hope you're doing well. I’ve reviewed your project requirements and see you are looking for a dataset of US patents in submarine cable technologies. You have no need to look any further; Zohaib is here to help you! My team has successfully completed over 50 similar projects. I will collect, clean, and geocode the patent data efficiently, ensuring high accuracy and usability within your budget. ➡️ Why Me? I can easily build your geocoded patent dataset as I have 5 years of experience in data collection, cleaning, and geocoding. My expertise includes working with patent databases, data analysis, and geospatial technologies. I also have a strong grip on Python and R, allowing me to create reproducible scripts for future use. ➡️ Let's have a quick chat to discuss your project in detail. I’d love to show you samples of my previous work. Looking forward to our conversation! ➡️ Skills & Experience: ✅ Data Collection ✅ Data Cleaning ✅ Geocoding ✅ Patent Analysis ✅ Python Programming ✅ R Programming ✅ SQL Database Management ✅ Data Visualization ✅ API Integration ✅ CSV & Excel Deliverables ✅ Geographic Information Systems (GIS) ✅ Name Disambiguation Waiting for your response! Best Regards, Zohaib
£350 GBP in 2 days
8.0
8.0

Hi — Elias here from Miami. I understand you're looking to build a geocoded USPTO patent dataset focused on submarine cable technologies. This project is crucial for extracting and analyzing patent trends effectively. What usually matters most here is managing and cleaning large datasets efficiently. The tricky part is ensuring accurate geocoding while maintaining data integrity throughout the extraction process. Given the complexities involved, a robust approach to data processing and storage will be essential. My approach would involve using Python with libraries like Pandas for data manipulation, ensuring scalability. I’d set up a pipeline to automate data extraction and cleaning, integrating with BigQuery for efficient data management. Future-proofing the system will allow for easy updates and additional data sources. I’ve worked on similar data-intensive projects where I streamlined workflows and improved reliability, ensuring actionable insights. A few questions to better understand the scope: Q1 – What specific data points are you looking to extract for the patents? Q2 – Are there any existing datasets or sources you plan to integrate with? Q3 – What is your timeline for this project, especially regarding data updates? Happy to discuss the details and suggest the best technical approach. Looking forward to hearing from you.
£500 GBP in 3 days
7.5
7.5

With a solid seven-year background in data science and medical data analysis, I am your go-to expert for this geospatial data extraction and cleaning project. My experience with different programming languages such as R and Python, including using tools like Pandas, will come in handy as I carry out this project. I hhave worked on various data projects before one of them being patent data extraction, for which I used BigQuery on Google Patents with great success. Regarding the geocoding of inventor addresses, I will leverage on powerful geocoding libraries like GeoPandas which can extract city-level lat/long while preserving the raw address components. I understand that there may be missing or malformed location data, but my strong problem-solving skills would enable me to handle these exceptions appropriately. Further, as you've highlighted, assignee name disambiguation is crucial in this project. To handle this task effectively, I would employ a combination of manual inspection and automated procedures that involve entity recognition and string matching to properly identify and collapse similar assignee names into one entity to ensure clean and coherent output. Lastly, also important is the project's timeline.I assure you timely completion without compromising quality. Hire me today for meticulous work output delivered with a high level of professionalism! Let me get to work producing results that not only meet your expectations but exceed them!"
£350 GBP in 3 days
7.1
7.1

Hi, I can build the patent dataset as a reproducible Python pipeline, starting with a few hundred records to validate search coverage, field extraction, location parsing, geocoding, and assignee normalization before scaling. I’d use PatentsView or USPTO/Google Patents data based on coverage and reproducibility, preserve raw location fields, geocode to city-level coordinates, document unresolved records, and deliver clean CSV, Excel, and source code. A few questions: * Will you provide the exact search terms and CPC classes, or should I help validate the initial query design? * Should inventor and assignee geocoding prioritize a specific provider such as Nominatim, Census Geocoder, or another source? * Do you want assignee normalization to include subsidiaries and parent companies, or only spelling/format variations of the same legal entity? Best regards, Muhammad Usman
£400 GBP in 4 days
6.8
6.8

Hello Client, I can build a reproducible USPTO patent pipeline using PatentsView/Google Patents, clean and disambiguate assignees, geocode locations to city-level coordinates, and deliver a validated sample before scaling.
£250 GBP in 2 days
6.7
6.7

Hello There! I’m Md Toriqul Islam, an experienced Python data engineer specializing in data collection, cleaning, entity matching, geocoding, and reproducible research pipelines. I understand you need a carefully validated sample of US patents related to submarine fibre-optic cable technologies, including patent metadata, inventor/assignee information, normalized entities, and city-level geocoded locations, before scaling to the full dataset. I have rich experience in Python, PatentsView/USPTO data, APIs, pandas, geospatial processing, and data validation. I am skilled in entity disambiguation, address normalization, geocoding, missing-data handling, and reproducible ETL pipelines. I can provide the sample as CSV/Excel plus a clean Python pipeline that can be rerun and extended. I’m available to complete the project within your 25 September 2026 deadline. I’m happy to share relevant data/geospatial project examples privately. Looking forward to hearing from you. Best regards, Md Toriqul Islam
£250 GBP in 4 days
6.4
6.4

I’d start with PatentsView as the primary source because it gives clean USPTO metadata, inventor/assignee locations, and reliable CPC search. I’d pull your sample cohort end-to-end, then validate the geocoding and assignee grouping before scaling to the full corpus. For geocoding, I’d use the Census Geocoder for US addresses and OpenStreetMap/Nominatim for international ones, falling back to state/country centroids only when a city-level match is impossible. Raw address fields stay intact in every row so you retain traceability. I’d flag unmatched and malformed records instead of silently dropping them. For assignee disambiguation, I’d normalize legal suffixes, punctuation, and case, then apply a hierarchical match: exact normalized match first, then string similarity against known parent names, and finally a manual review of the top residual entities. This gets Corning variants collapsed into one entity without merging unrelated companies. The deliverable would be a reproducible Python pipeline that outputs your sample processed, clean, and geocoded as CSV plus Excel, ready for you to inspect before I run the full extraction.
£250 GBP in 7 days
6.0
6.0

Patent data is structured enough to extract reliably and messy enough to punish naive scraping — claims, classifications and citations each need their own handling. - Python extraction from your source (EPO/USPTO bulk data or the site you specify), respecting rate limits - Parsing into clean tabular data: numbers, dates, assignees, IPC/CPC classes, claims - Analytics layer: trends by assignee/class/time, delivered as Excel/CSV plus charts Proof: I build document-processing and data-extraction pipelines in production for industrial clients, running unattended. Which source and roughly what volume — and what question should the analytics answer (competitor landscape, whitespace, filings over time)? Happy to start with a sample batch. Martin
£349 GBP in 3 days
6.0
6.0

I’ll build a clean, reproducible geocoded USPTO patent dataset for submarine fibre-optic cable technologies, end-to-end from sourcing to analytics-ready CSV/Excel and scripts. Approach: pull records from a justified source (PatentsView, Google Patents/BigQuery, or USPTO bulk), filter by the CPC/classes and search terms you provide, then standardize fields (patent number, title, abstract, filing/grant dates, assignees, inventors). I’ll disambiguate assignee names by normalization plus rule-based entity merging (e.g., Corning Incorporated/Inc variants) and preserve provenance. Geocoding: inventor and assignee locations will be parsed into structured components (city/state/country) and converted to city-level lat/long using a documented, repeatable geocoding workflow. Missing or malformed locations will be handled with explicit fallbacks and audit fields so you can quantify coverage before scaling. Deliverables: a fully processed “few hundred patents” sample to validate the pipeline, then scalable processing to the full corpus; CSV + Excel outputs and a Python/R script to rerun and extend. Cheerfully positive collaboration focused on getting the pipeline right, then scaling cleanly.
£250 GBP in 3 days
5.4
5.4

The search terms and CPC classes can be interpreted in multiple ways, leading to incomplete or irrelevant patent sets. I will pull patent records from USPTO bulk data, specifically using their patent assignment data and classification information, also looking at publicly available Google Patents BigQuery datasets for completeness, and I will focus on the primary technology area of submarine fibre-optic cable systems, also including related optical and materials technologies as specified. I will build the dataset using Python, writing scripts to query these sources and then process the raw data into a clean format, so that duplicate patents are removed and each entry is standardized. The geocoding will be done using the inventor addresses from the patent data, mapping these to their respective US counties and then to commuting zones using publicly available shapefiles and a Python library like `geopandas`. The part that breaks on this job is when inventor addresses are incomplete or ambiguous, which can lead to geocoding failures. I will build in a step to flag these records and perform targeted manual review and cross-referencing against other patent data sources to assign the most accurate geocodes possible so that the geographic analysis is sound. Your attached paper uses commuting zones as the geographic unit, so will you need the patent data mapped to commuting zones or will county-level data suffice for your analysis? I have 8 reviews on here, everything delivered on time and on the agreed price so far, plus Preferred Freelancer status. Once you provide the search terms and CPC classes, I will send back a sample of the first 50 extracted and geocoded patents so you can see the data structure and quality.
£578 GBP in 21 days
5.3
5.3

Hello!, This is James from Hollywood... I read your post carefully, and this is exactly the kind of data project where the details matter. You need more than a scrape, you need a clean, geocoded USPTO patent dataset focused on submarine cable technologies, with proper cleaning and a structure you can actually analyze in Excel or BigQuery. My approach would be simple and effective: 1. Identify the right USPTO sources and patent filters for the scope 2. Extract the key fields with Python/Pandas 3. Normalize assignees, inventors, dates, citations, and locations 4. Geocode the location data and remove duplicates/inconsistencies 5. Deliver a clean CSV/Excel plus a BigQuery-friendly dataset ready for analysis I’m very strict about data quality, because in patent landscape work one bad field can throw off the whole analysis. I’d rather get the structure right from the start than hand over something that only looks finished. A few quick questions: - Do you want granted patents only, or also published applications? - Should geocoding cover inventor location, assignee HQ, or both? - Do you already have a preferred data source, or should I build the workflow from scratch? Relevant work I’ve done includes patent and research data pipelines, a market intelligence scraper for a niche manufacturing dataset, and a geocoded business dataset used in BigQuery analytics. If useful, I can also suggest the schema before starting so you know the dataset will be usable from day one.
£600 GBP in 3 days
5.5
5.5

Hi, the hardest part here is not collecting patent records, but preserving a clean chain from the original source data through assignee disambiguation and geocoding so the final dataset is reliable enough for geographic analysis. My approach would be to first build the pipeline on a few hundred patents and validate every field before scaling. I’d select the source based on coverage of your search terms, CPC classes and inventor/assignee location data, while keeping the extraction logic reproducible in Python. For geocoding, I’d preserve city, state and country as separate raw fields, normalize malformed locations before geocoding, and keep unmatched or ambiguous records clearly flagged rather than silently assigning uncertain coordinates. The geocoded output would include validation information so later mapping to commuting zones remains auditable. For assignee names, I’d combine normalization rules with a controlled matching process: case and punctuation normalization, removal of obvious legal suffix variation, then review of uncertain matches to avoid incorrectly merging different entities. One important risk is over-aggressive entity matching. Two distinct assignees can appear similar, so I’d keep the original assignee string alongside the standardized entity. A few questions: Q1: Are the search terms and CPC classes already finalized? Q2: Should inventor and assignee locations be geocoded globally or only within the US? Juan Pablo
£500 GBP in 7 days
5.4
5.4

I am interested in assisting with your Patent Data Extraction and Analytics project, bringing strong attention to detail and a structured approach to handling complex patent information. I can efficiently extract relevant data from patent documents and databases, including patent numbers, titles, applicants, inventors, filing dates, classifications, citations, claims, and other required fields while maintaining accuracy and consistency. My approach will include organizing the extracted information into clean, structured datasets suitable for analysis. I can perform data cleaning, validation, categorization, and deduplication, as well as analyze patent trends, technology areas, competitors, filing activity, and citation patterns. I will ensure that the data is presented in a clear format that supports meaningful business and intellectual-property insights. I am committed to delivering accurate, well-organized, and reliable results within the agreed timeline. I can adapt the extraction process to your preferred sources, formats, and analytical requirements and provide regular progress updates throughout the project. I would be glad to discuss your specific patent data requirements and demonstrate how I can contribute effectively to the project.
£250 GBP in 7 days
5.4
5.4

Warm greetings, we can build a clean, reproducible geocoded USPTO patent dataset and validate a few hundred records end-to-end before scaling. We are a team of 62 professionals with over 9 years of experience in Python, Pandas, BigQuery, data extraction, cleaning, and geospatial datasets. Here's how we can help: * Pull patents using your search terms/CPC classes via PatentsView or Google Patents BigQuery * Extract patent, assignee, inventor, filing/grant, and location fields * Normalize assignee variants using rule-based matching plus fuzzy/entity resolution * Geocode city/state/country while preserving raw fields and flagging malformed/missing data * Deliver CSV, Excel, and fully reproducible Python scripts For malformed locations, we would standardize and validate first, then geocode only reliable records and retain confidence/status fields for review. Do you already have the search terms/CPC list and the attached reference paper available for the sample phase?
£500 GBP in 7 days
5.4
5.4

Your geocoding pipeline will fail if you treat USPTO inventor addresses as clean input — they're inconsistent across decades and contain obsolete city names that won't resolve without historical gazetteers. This will corrupt your commuting-zone mapping. Quick questions - are you planning to map patents filed before 1990, when address standardization was poor? And do you need FIPS code linkage for commuting zones, or will lat/long suffice for your spatial join? Here is the architectural approach: - PATENTSVIEW + BIGQUERY: Pull from PatentsView API for post-2000 records and BigQuery for historical bulk data, then merge on patent number to fill gaps in inventor location fields. - GEOCODING PIPELINE: Use Google Geocoding API with fallback to OpenCage for ambiguous addresses, logging confidence scores per record so you can flag low-quality matches before commuting-zone assignment. - ASSIGNEE DISAMBIGUATION: Build a fuzzy-matching script using Python's RapidFuzz library to cluster name variants, then manually review clusters above 80% similarity to catch edge cases like subsidiaries. I've built similar patent datasets for a cleantech research group that required CPC-class filtering and multi-decade geocoding with FIPS linkage. Let's do a 200-patent sample run first so you can validate the output format before I scale to the full corpus.
£450 GBP in 21 days
5.4
5.4

Hi, Your sample-first approach is exactly how I would handle this. I would build the extraction and geocoding pipeline on a few hundred patents first, validate the resulting patent, inventor, assignee and location records with you, then use the approved pipeline for the full corpus. I can implement this in Python using Pandas with either Google Patents/BigQuery or PatentsView depending on your supplied CPC classes and search methodology. For locations, I would preserve the original city/state/country fields, create standardized versions, geocode to city-level coordinates, cache matches, and explicitly flag unresolved or ambiguous records rather than silently guessing. For assignees, I would normalize punctuation, capitalization and corporate suffixes first, then apply controlled fuzzy/entity matching so variants such as Corning Inc. and Corning Incorporated resolve to a canonical entity while preserving the original value. Deliverables will include the validated CSV/Excel datasets, reproducible Python pipeline, configuration/search inputs, and documentation so the corpus can be regenerated or extended later.
£395 GBP in 7 days
4.9
4.9

Hi I can help build a clean, reproducible USPTO patent dataset for submarine fibre-optic cable technologies, including data extraction, cleaning, entity normalization, and geographic mapping. I have experience with Python, data pipelines, web/API extraction, Pandas, geospatial processing, and large structured datasets. I can build the workflow using sources such as PatentsView, USPTO bulk data, or Google Patents BigQuery depending on coverage and reproducibility requirements. My approach would start with a few hundred patents as a complete sample pipeline, including patent metadata extraction, inventor/assignee cleaning, location parsing, city-level geocoding, validation, and export to CSV/Excel. After confirming accuracy, I can scale the process to the full patent corpus. For geocoding, I would preserve raw address fields, normalize locations, use reliable geocoding services, and handle missing or ambiguous addresses through validation rules and documented exceptions. For assignee disambiguation, I would apply normalization techniques, aliases, and matching rules to consolidate organization names consistently. I will also provide the Python/R scripts and documentation so the dataset can be reproduced and extended for future research. Best, Justin
£500 GBP in 7 days
5.0
5.0

Your project to geocode USPTO patents for submarine cable technologies aligns perfectly with my experience in building structured, geographically-aware datasets for economic analysis. I've successfully executed similar projects involving large-scale patent data extraction and enrichment, including the development of geocoded innovation maps for emerging tech sectors that mirrored the analytical goals of your attached paper. My understanding of patent classification codes and the nuances of extracting relevant technical details from USPTO documents ensures a high-quality foundation for your research. My approach will involve leveraging Python with libraries like `pandas` for data manipulation and `requests` for accessing patent data APIs. I'll employ a multi-stage process: first, programmatically querying the USPTO database using relevant CPC codes and keywords to gather patent metadata. Second, I'll perform robust data cleaning, standardizing patent assignee names and addresses. Finally, I'll utilize geocoding services (e.g., Google Geocoding API or an open-source alternative like Nominatim) to convert cleaned address data into precise latitude and longitude coordinates, ensuring accuracy for your spatial analysis. To ensure optimal alignment with your analytical needs, could you clarify if you have a preferred geocoding service or any specific data fields beyond assignee location that are critical for your mapping? I'm confident I can deliver a precise and actionable geocoded dataset. I'm available for a brief call to discuss the specifics and how I can best contribute to your academic project.
£578 GBP in 21 days
4.6
4.6

Hi, this is exactly the kind of dataset build I can help with: pulling USPTO patent records, cleaning assignee names, and producing a geocoded, analysis-ready file for your submarine cable technology study. I’ve worked on reproducible data pipelines that combine patent metadata, entity cleaning, and location standardization in Python and pandas, with outputs designed for research use. For the sample phase, I’d first assemble a few hundred patents from the source you prefer, validate the fields end to end, and keep raw location components separate from the geocoded city-level fields. For geocoding, I’d use a consistent, auditable process with fallbacks for missing or malformed addresses, then flag uncertain matches rather than forcing them. For assignee disambiguation, I’d normalize names, strip punctuation, compare variants, and build a reference table so the logic is reproducible. Happy to discuss the sample scope and workflow. Best regards, Gabriel
£250 GBP in 7 days
4.1
4.1

Hi there, I’d be excited to build your USPTO patent dataset with an accuracy-first, reproducible pipeline. I understand the goal is not simply collecting patents, but producing clean, geocoded data suitable for quantitative analysis of where submarine cable innovation occurs. I can build this in Python using Pandas and the most suitable source between PatentsView, Google Patents BigQuery, or USPTO bulk data. I’ll extract patent, inventor, assignee, date, and location fields, preserve raw address components, and normalize the dataset for analysis. For geocoding, I’ll convert inventor and assignee locations to city-level coordinates while retaining the original fields. Missing or malformed locations will be flagged rather than silently discarded. I’ll also normalize assignee names and carefully consolidate legitimate variants without aggressive false matches. I’ll begin with the requested few-hundred-patent sample, validate the pipeline with you, then scale it to the full corpus. Delivery will include CSV, Excel, and a documented Python script that you can rerun and extend. I think you want research-grade, transparent data rather than a one-off spreadsheet, and I’ll build the workflow accordingly. Looking forward to working with you. Thanks
£300 GBP in 5 days
4.0
4.0

Oxford, United Kingdom
Payment method verified
Member since Aug 24, 2026
₹600-800 INR
₹12500-37500 INR
₹37500-75000 INR
$8-15 USD / hour
$250-750 USD
$30-250 USD
₹1500-12500 INR
$30-250 USD
$250-750 USD
₹1500-12500 INR
$30-250 USD
$250-750 USD
₹1500-12500 INR
₹37500-75000 INR
$30-250 USD
₹600-1500 INR
₹1250-2500 INR / hour
$25-50 USD / hour
$30-250 USD
$30-250 USD