Filter

My recent searches
Filter by:
Budget
to
to
to
Type
Skills
Languages
    Job State
    1,077 pyspark jobs found

    Azure Data Factory + Azure SQL Expert Needed to Debug and Implement Data Processing Change Description I'm looking for an experienced Azure Data Engineer with strong knowledge of Azure Data F...Data Flow JSON Stored Procedures SQL View definitions Ticket description Access during a screen-sharing session (if required) Deliverables Identify the exact implementation points. Implement the required changes. Explain every modification. Provide testing steps. Help validate the solution. Required Skills Azure Data Factory Azure SQL SQL Server Stored Procedures Mapping Data Flows Databricks / PySpark Azure Data Engineering Debugging existing enterprise ETL solutions Note The work will be done through screen sharing, and I cannot share confidential source code or business data outsi...

    $33 Average bid
    $33 Avg Bid
    12 bids

    ...materialized masked views. Experience with schema drift monitoring and metadata-driven ETL frameworks. Strong understanding of data governance, data security, and regulatory compliance. Experience with ETL/ELT development and enterprise data warehousing. Familiarity with Git, CI/CD pipelines, and Agile methodologies. --- Preferred Skills Azure Data Lake Storage (ADLS) Azure Synapse Analytics PySpark Python Azure DevOps PowerShell Microsoft Purview or similar data governance tools --- Qualifications Bachelor's or Master's degree in Computer Science, Information Technology, or a related field. Strong analytical, troubleshooting, and communication skills. Experience working with global teams and US-based clients is highly preferred. --- Engagement...

    $5 / hr Average bid
    $5 / hr Avg Bid
    23 bids

    ...Upcoming Newly Added Topics --- ## 4. Interview Question Library This is the core feature of the website. Questions should be organized topic-wise. Example categories: ### SQL * Basic SQL * Intermediate SQL * Advanced SQL * Window Functions * CTE * Joins * Performance Optimization --- ### Python * Basics * OOP * Collections * Iterators * Generators * Interview Coding Questions --- ### PySpark * DataFrames * RDD * Spark Architecture * Transformations * Actions * Partitioning * Performance Optimization * Spark UI * AQE * Shuffle --- ### Azure * Azure Data Factory * Azure Data Lake * Azure Synapse * Azure Functions * Logic Apps * Managed Identity * Networking --- ### Databricks * Cluster * Unity Catalog * Workflows * Delta Live Tables * Auto Loader * Photon * Serv...

    $318 Average bid
    $318 Avg Bid
    77 bids

    I need an end-to-end ETL pipeline that moves data from our on-premise databases into Google Cloud, transforms it, and lands it cleanly in BigQuery. The core stack must be Airflow for orchestration, Dataproc (running PySpark) for heavy transformations, and native BigQuery SQL for final modelling and reporting layers. You will design and implement: • Secure ingestion from the on-prem source into GCS staging • Airflow DAGs that trigger Dataproc jobs, handle retries, logging and alerting • PySpark transformation scripts on Dataproc, tuned for performance and cost • BigQuery SQL models that expose the refined tables • Parameterised configuration so environments can be promoted from dev to prod without code changes Acceptance criteria: &nd...

    $495 Average bid
    $495 Avg Bid
    11 bids

    ...Downstream jobs remain healthy, so the issue is isolated to the first step of the pipeline that pulls structured data into Databricks before any transformations begin. I need you to jump into the existing notebooks and jobs, trace the root cause, and deliver a clean, repeatable fix. The environment runs on Databricks Runtime 12.x with Auto Loader feeding into Delta tables, so familiarity with Spark, PySpark, SQL, and cluster-level configs is essential. I will grant you workspace access and point you to the failing job runs and relevant logs. Acceptance criteria: • Ingestion job completes successfully three runs in a row with identical input files. • No data loss or duplication in the target Delta tables (row counts match source). • A concise summary of what c...

    $21 Average bid
    $21 Avg Bid
    30 bids

    ...CSV/SQL sources and processes it in Spark notebooks. The jobs run fine until the transformation stage, where they suddenly throw runtime errors and the cluster crashes. I need someone who can jump in quickly, reproduce the failure, trace the exact root cause, and deliver a rock-solid fix without disrupting the live data feed. You should be fully comfortable with Databricks Runtime, Spark (Scala or PySpark), Delta Lake tables, cluster configuration, and Azure storage connectors, as that’s the stack in play. If a configuration tweak, code refactor, or resource-scaling change is required, implement it, document the change, and prove stability with reruns under production-level load. Deliverables (all required): • Root-cause analysis summary with error logs highlighted...

    $12 Average bid
    $12 Avg Bid
    9 bids

    Python, SQL, ETL, PySpark, Spark SQL, AWS EMR, AWS Lambda, AWS Step Functions, Amazon S3 (Data Lake – Raw & Processed Zones), AWS CloudWatch, AWS SNS, Pandas, Excel, Veeva CRM Project Overview Designed and implemented an end-to-end AWS-based data engineering pipeline to bifurcate, process, and deliver pharmaceutical sales, HCP, call activity, territory, and marketing data for Europe (EU) and Russia (RU) regions into Veeva CRM. The solution automated data ingestion from external APIs, validated and transformed high-volume datasets using Spark on EMR, and enforced multi-layer data quality checks based on business rules. Final curated datasets were delivered to Veeva CRM to support daily call planning, HCP targeting, territory alignment, and field sales insights, enabling acc...

    $10 / hr Average bid
    $10 / hr Avg Bid
    15 bids

    ...screen recording Hands-on lab walkthroughs for each module Supporting slides, code samples, and datasets Topics to cover (full syllabus provided) Python & SQL foundations, Linux & Git PostgreSQL, data modeling, star schemas, SCDs Cloud warehousing — BigQuery & Snowflake (partitioning, cost tuning) ELT pipelines with dbt (models, tests, snapshots, lineage) Workflow orchestration with Apache Airflow PySpark for big data + real-time streaming with Kafka & Spark Structured Streaming Docker, CI/CD (GitHub Actions), Terraform basics, Azure Data Factory End-to-end capstone pipeline (ingest → transform → serve) Requirements 3+ years of hands-on data engineering experience Clear spoken English and good audio/screen recording quality Ability to teach begin...

    $95 Average bid
    $95 Avg Bid
    12 bids

    ...years of curated historical performance data. The goal is to turn those two sources into accurate failure-risk forecasts and actionable maintenance recommendations. You will design and implement the full workflow—data ingestion, cleansing, feature engineering, model training, and deployment—inside a Python-based stack (pandas, scikit-learn or similar; feel free to propose alternatives such as PySpark or TensorFlow if they offer clear advantages). Real-time scoring must run fast enough to support control-room dashboards, and the results should be exposed through an API the rest of the system can call. Key deliverables • Data pipeline that pulls JSON sensor streams and historical records into a unified store • Predictive models with documented performanc...

    $295 Average bid
    $295 Avg Bid
    70 bids

    ...domain workflows. • Deploy scalable, resilient agent pipelines with monitoring and evaluation. GenAI Application Engineering • Develop GenAI applications using models like GPT, Gemini, and LLaMA. • Implement RAG, vector search, prompt orchestration, and model evaluation. • Partner with data scientists to productionize POCs. Data & Platform Engineering • Build distributed data pipelines (Python, PySpark). • Develop APIs, SDKs, and integration layers for AI-powered applications. • Optimize systems for performance and scalability across cloud/hybrid environments. MLOps / LLMOps • Contribute to CI/CD workflows for AI models—deployment, testing, monitoring. • Implement governance, guardrails, and reusable GenAI frameworks. Collabo...

    $597 Average bid
    $597 Avg Bid
    70 bids

    ...handling, logging, and file validation checks within the shell pipeline to ensure zero data loss during ingestion. Data Cleaning & Preprocessing (Python): * Utilized Python (Pandas/NumPy) to handle raw, messy source data. Resolved structural data issues, treated missing values, eliminated duplicates, and standardized data formats for downstream consumption. Scalable ETL Processing (PySpark): * Leveraged PySpark to process and transform large-scale datasets efficiently, optimizing partition strategies to reduce execution time. Executed complex business logic, aggregations, and data joins across massive distributed data frames. Hybrid Data Storage (PL/SQL & MongoDB): Relational (PL/SQL): Designed relational schemas, wrote optimized stored procedures, triggers, and co...

    $246 Average bid
    $246 Avg Bid
    26 bids

    ...customer-facing features - Support ML/MLOps workflows including model training and deployment - Implement monitoring, error handling, and automated recovery systems - Ensure data quality, validation, and anomaly detection - Design reusable, scalable data schemas and APIs - Contribute to CI/CD pipelines with automated testing and deployments Required Skills: - Strong Python (5+ years), SQL, Pandas/PySpark - AWS expertise (Lambda, Glue, S3, Athena, DynamoDB, ECS Fargate) - Experience with data lakes and serverless architectures - Backend/API development (REST, microservices, event-driven systems) - Working knowledge of Node.js or willingness to learn - Exposure to ML/MLOps (SageMaker, model lifecycle) - Experience with large-scale data processing and distributed systems - Strong ...

    $648 Average bid
    $648 Avg Bid
    26 bids

    ...enforce engineering best practices --- Required Skills 7–10 years in Data Engineering (3+ years in lead role) Strong in SQL, Python, and Spark (PySpark/Scala) Experience with Snowflake, Databricks, BigQuery, or Redshift Solid understanding of data modeling (Kimball / Data Vault) Hands-on with streaming systems (Kafka / Flink) Familiarity with Terraform, CI/CD, and Cloud (AWS/GCP/Azure) --- Good to Have Experience with Feature Stores (Feast, Tecton) Knowledge of Data Mesh / CDC tools (Fivetran, Airbyte) Exposure to Graph or Vector Databases Open-source contributions or advanced degree --- Tech Stack Python, SQL, PySpark, Scala | Snowflake, Databricks, BigQuery Airflow, dbt | Kafka | AWS/GCP | Docker, Kubernetes, Terraform --- Interested candi...

    $1280 Average bid
    $1280 Avg Bid
    12 bids

    ...so we can swap to JSON or Parquet later without touching business logic. • Store raw and feature-engineered data in partitioned Parquet inside S3, queried locally through DuckDB/Polars and cached in Redis. • Train and serve forecasts with LightGBM and Croston/TSB, orchestrated by Airflow. • Ship a single-node pipeline that is production-ready yet cleanly abstracted for a 2–4 week lift to PySpark on EMR Serverless when volumes grow. • Expose the pipeline through createServerFn so the existing React storefront can request forecasts in real time. What must be fully exercised by end-to-end tests (Playwright + Vitest + pytest): • Data ingestion & processing • Machine-learning prediction paths • Data storage & retrieval ...

    $661 Average bid
    $661 Avg Bid
    164 bids

    ...SageMaker) - Implement monitoring, error handling & ensure high system reliability (99.9% uptime) - Build data validation, quality checks & anomaly detection systems - Design systems for backfills, reprocessing & consistency - Maintain data contracts, schema versioning & CI/CD pipelines --- Required Skills - 3–7+ years of software/data engineering experience - Strong Python (5+ yrs), SQL, Pandas/PySpark - Hands-on AWS (Lambda, Glue, S3, Athena, DynamoDB, ECS, Step Functions) - Experience with REST APIs, microservices, event-driven architecture - Knowledge of ML/MLOps (SageMaker, model lifecycle) - Exposure to Node.js (or willingness to learn) - Experience with large-scale data processing & distributed systems - Strong focus on testing, CI/CD, monitor...

    $619 Average bid
    $619 Avg Bid
    35 bids

    ...end-to-end on AWS, coded in PySpark and designed for production reliability. The pipeline will ingest data from three sources—databases, APIs, and file systems—then standardise and load it into an analytics-ready destination. Source files arrive in a mix of CSV, JSON, and Parquet, so the job must include automatic format detection, schema inference, and efficient column-wise writes. Beyond raw transformation, I want solid engineering practices: parameter-driven jobs, modular Spark code, unit tests, logging, alerting, and retry logic. Leveraging AWS native services such as Glue, EMR, Lambda, and S3 is expected, but I’m open to other AWS components if they shorten development time or lower cost. Candidates must have expertise in data engineering. Deliverables ...

    $26 / hr Average bid
    $26 / hr Avg Bid
    24 bids

    ...Responsibilities: - Develop ML/statistical models (DID, Synthetic Control, A/B Testing) in Python - Build and integrate FastAPI-based services - Design large-scale data pipelines using PySpark, Delta Lake, and Azure Data Lake - Optimize Spark jobs (memory, partitioning, performance tuning) - Work with Databricks for job orchestration and data workflows - Containerize and deploy applications using Docker & Kubernetes - Ensure code quality with testing and CI/CD pipelines - Collaborate with data science and product teams --- Must Have Skills: - Python (3.9+), Pandas, NumPy, Scikit-learn, SciPy - Strong PySpark & Spark Internals (OOM handling, tuning, optimization) - Databricks (clusters, workflows, Delta Lake) - Causal Inference: A/B Testing, DID, Hypothesis Test...

    $584 Average bid
    $584 Avg Bid
    31 bids

    ...Unity Catalog enabled, and I need a seasoned modeller who can translate business requirements into robust Star and Snowflake schemas, then bring them to life with PySpark and advanced SQL. You will refine our Medallion architecture (Bronze → Silver → Gold), implement both Type 1 and Type 2 SCD strategies, and tune the pipelines for speed through smart partitioning and other optimisation techniques. The datasets involved are large, structured and semi-structured, so hands-on experience handling such volumes in Databricks is essential. Key deliverables • Logical and physical data models documented and version-controlled • PySpark notebooks / SQL scripts that create the Star and Snowflake tables in Delta Lake under Unity Catalog governance • P...

    $11 / hr Average bid
    $11 / hr Avg Bid
    13 bids

    ...should be an active contributor with the ability to handle complex data challenges independently within an existing architecture. Key Responsibilities: Design, develop, and maintain data pipelines using PySpark Work on data ingestion, transformation, and optimisation for large-scale datasets Handle real-world data challenges such as API inconsistencies, schema drift, and incremental load failures Ensure data quality, reliability, and performance across pipelines Collaborate with cross-functional teams to deliver data-driven solutions Required Skills: Strong hands-on experience in PySpark and data engineering Proven experience in handling production-level data issues and debugging Solid understanding of data modelling, ETL/ELT processes, and data pipelines Ability to wo...

    $13 / hr Average bid
    $13 / hr Avg Bid
    8 bids

    ...should be an active contributor with the ability to handle complex data challenges independently within an existing architecture. Key Responsibilities: Design, develop, and maintain data pipelines using PySpark Work on data ingestion, transformation, and optimisation for large-scale datasets Handle real-world data challenges such as API inconsistencies, schema drift, and incremental load failures Ensure data quality, reliability, and performance across pipelines Collaborate with cross-functional teams to deliver data-driven solutions Required Skills: Strong hands-on experience in PySpark and data engineering Proven experience in handling production-level data issues and debugging Solid understanding of data modelling, ETL/ELT processes, and data pipelines Ability to wo...

    $13 / hr Average bid
    $13 / hr Avg Bid
    6 bids

    I need a Microsoft Fabric notebook written in PySpark that can call a suitable commodities-exchange API and pull the most recent futures data on a couple select futures products. The solution should be fully runnable inside Fabric. Key points I have set: • Refresh cadence: once a month, so include a simple scheduling example (Fabric pipeline or a cron-style note is fine). This data will be pulled from a Chinese exchange, so I want a Chinese speaking freelancer only!! Deliverables 1. The .ipynb (or .notebook) file ready to import into Fabric 2. A quick test run showing one successful fetch and a tidy DataFrame with the fields timestamp, contract, price, and volume Acceptance will be based on the notebook executing end-to-end without manual edits (apart from entering an...

    $30 Average bid
    $30 Avg Bid
    17 bids

    ...and transform it in real time, feed it to a set of AI services, then serve the insights back to users through an intuitive dashboard. Your day-to-day work will touch three key areas: • Data collection – build reliable connectors, handle auth flows, schedule recurring pulls, and maintain error logging. • Data processing – design ETL pipelines, implement transformation logic in Python (Pandas, PySpark or similar), and ensure everything is containerised for smooth deployment. • Data visualization – wire processed datasets into the React front-end, craft reusable chart components (D3, , or your preferred library), and optimise for performance. Acceptance criteria 1. End-to-end pipeline runs with a one-command deploy (Docker / docker-compose or...

    $100 Average bid
    $100 Avg Bid
    36 bids

    Data cleaning using SQL/Python (need to figure out) and export in Excel. The client have used this in Pyspark environment. we can have a discussion later ont the details.

    $9 Average bid
    $9 Avg Bid
    1 bids

    We are looking for an experienced Palantir Foundry Developer to support data and AI use cases. Scope of Work: * Build and maintain Foundry data pipelines (Pipeline B...support data and AI use cases. Scope of Work: * Build and maintain Foundry data pipelines (Pipeline Builder, Transforms) * Work with Ontology (object types, link types, data modeling) * Develop Workshop applications for business users * Implement AIP Logic workflows and basic agent integrations * Write production-quality Python, SQL, and PySpark code Requirements: * Hands-on experience with Palantir Foundry (mandatory) * Strong skills in Python, SQL, and PySpark * Experience with Ontology, Pipelines, and Workshop * Basic understanding of AIP (preferred) Project Details: * Budget: ₹45,000+(Negotiable) * ...

    $537 Average bid
    $537 Avg Bid
    21 bids

    Results-driven Senior Data Analyst with 8+ years delivering enterprise data solutions across Banking, Financial Services, and Healthcare. Specialized in end-to-end Data Warehouse design, ETL pipeline development, and BI reporting. Core Expertise: Snowflake, Azure Data Factory, Azure Synapse, Delta Lake, PySpark, SSIS, T-SQL, PL/SQL, Oracle, MySQL, Power BI, DAX, Tableau, SSRS, Star/Snowflake Schema, EDW Design, Data Mart Development, Data Lineage, Gap Analysis. Compliance & Governance: HIPAA, GDPR, SOX Audit Controls, Data Quality Frameworks, Data Governance Policies. Business Analysis: BRD/FRD Writing, JAD Facilitation, UML Diagrams, Stakeholder Management, Agile, Waterfall. Certifications: Microsoft Azure Fundamentals, Salesforce Administrator, Salesforce Platform Developer I....

    $150 Average bid
    $150 Avg Bid
    23 bids

    Responsible for designing and implementing large-scale data migration and ingestion pipelines to move high-volume data from diverse sources into cloud platforms. Sources include HDFS, relational databases such as MySQL and PostgreSQL, and real-time streaming systems like Kafka. Develop and maintain robust data pipelines using PySpark, ensuring efficient processing of batch and streaming data. Implement automated scheduling mechanisms to orchestrate data workflows on daily and monthly intervals, ensuring reliability and timely data availability. Optimize data ingestion and storage through advanced performance tuning, partitioning, and compaction strategies to handle large-scale datasets efficiently. Ensure data quality, consistency, and fault tolerance across all pipelines. Deploy...

    $10 Average bid
    $10 Avg Bid
    1 bids

    ...Experience Required: 5+ Years (Data Engineering), 3+ Years (Databricks) Note: Budget is fixed. Please do not apply if you are looking to negotiate. Key Responsibilities Develop and optimize data pipelines using Databricks, PySpark, and Spark SQL Design and implement Delta Lake architecture (Bronze / Silver / Gold layers) Work on Lakehouse architecture and manage Unity Catalog Apply DataOps practices for scalable and reliable data workflows Optimize Spark jobs for performance and cost efficiency Required Skills Strong hands-on experience with Databricks Proficiency in PySpark and Spark SQL Experience with Delta Lake and Lakehouse architecture Knowledge of data pipeline design and optimization Understanding of DataOps and data governance Nice to Have Experience with Azur...

    $625 - $729
    Sealed NDA
    $625 - $729
    11 bids

    ...database containing nested JSON / key-value blobs. • Goal: parse, normalize, and flatten these blobs into well-defined columns while preserving relationships and lineage. • Scale: millions of rows, so solutions that leverage Spark, Hadoop, BigQuery, Snowflake, or well-tuned SQL/Python pipelines are welcome—as long as they remain maintainable. Deliverables 1. Transformation code (Python, PySpark, SQL, or Scala) with clear comments. 2. A runnable job definition or workflow file (Airflow DAG, Spark submit script, dbt model, etc.) that shows how to execute the pipeline end-to-end. 3. Simple README explaining prerequisites, run steps, and how new fields should be added in future. Acceptance criteria • Pipeline processes at least 10 GB of source data ...

    $148 Average bid
    $148 Avg Bid
    6 bids

    ...validation rules, automated tests, and observable metrics baked in from day one—Great Expectations, Delta Live Tables expectations, or comparable frameworks are welcome, as long as quality gates are visible in the monitoring layer. Scope to cover: • Architecture design diagram with clear component rationale (Azure Data Lake, Databricks, Delta, Unity Catalog, etc.). • Reproducible code (Python / PySpark, notebooks or repos) with CI/CD instructions. • Ingestion pipelines (batch or streaming), curated layers, and serving tier (SQL endpoints, Power BI, or dashboards of your choice). • Integrated monitoring, alerting, and cost-aware observability using native Azure tools or open-source add-ons. • End-to-end test suite: unit, integration, and data qualit...

    $11 / hr Average bid
    $11 / hr Avg Bid
    19 bids

    ...validation rules, automated tests, and observable metrics baked in from day one—Great Expectations, Delta Live Tables expectations, or comparable frameworks are welcome, as long as quality gates are visible in the monitoring layer. Scope to cover: • Architecture design diagram with clear component rationale (Azure Data Lake, Databricks, Delta, Unity Catalog, etc.). • Reproducible code (Python / PySpark, notebooks or repos) with CI/CD instructions. • Ingestion pipelines (batch or streaming), curated layers, and serving tier (SQL endpoints, Power BI, or dashboards of your choice). • Integrated monitoring, alerting, and cost-aware observability using native Azure tools or open-source add-ons. • End-to-end test suite: unit, integration, and data qualit...

    $546 Average bid
    $546 Avg Bid
    58 bids

    ...transformation, and optimisation. • Hands-on experience working within Databricks, including notebooks, workflows, and job execution. • Proven experience using Power BI for report and dashboard development, including data modelling, DAX, Power Query, and visualisation design. • Experience building and maintaining data pipelines, ideally within Azure environments. • Experience using Python (e.g. PySpark) within Databricks environments is advantageous. • Understanding of data modelling concepts, including fact and dimension structures. • Familiarity with Azure Data Factory or similar orchestration tools, with Insight Factory advantageous. • Working knowledge of DevOps practices, including version control, repository management, and...

    $1370 Average bid
    $1370 Avg Bid
    29 bids

    Job Title: Data Engineer (Databricks and AWS) Duration: 2 Hours Budget: ₹22,000 – ₹26,000 (based on screening) Tech Stack: Databricks, Python, PySpark, AWS, SQL, Git Job Description: We are looking for an experienced Data Engineer to provide short-term support. The role involves working on data pipelines, transformations, and analytics using Databricks and AWS. Responsibilities: Develop and optimize data pipelines using Databricks, PySpark, and Python Work with AWS services and SQL-based data processing Manage code and versioning using Git Troubleshoot and optimize data workflows Requirements: Strong hands-on experience with Databricks, PySpark, and Python Good knowledge of AWS data services and SQL Experience with Git and collaborative development Ability to del...

    $295 Average bid
    $295 Avg Bid
    16 bids

    ...engineering experiences with various aws services Experience building end-to-end data pipelines (schema discovery, ingestion, transformation, orchestration, monitoring) Experience working with relational databases like Oracle, MySQL, and SQL Server etc Experience with data ingestion from on-prem systems to cloud Experience with streaming platforms like Kafka or AWS Kinesis Strong skills in Python, PySpark, SQL, and Terraform...

    $1132 Average bid
    $1132 Avg Bid
    157 bids

    ...the next round of hiring I want an accomplished Senior Data Engineer to sit in on our technical interviews for roughly two hours each day. The role is purely evaluative: you will craft probing questions, join live video calls, and quickly score each candidate’s depth of knowledge across Python, Scala and SQL. Our stack centres on Azure and Databricks, so practical insight into large-scale Spark/PySpark jobs, data-model design, ETL orchestration and cloud performance tuning is essential. Candidates frequently discuss streaming, optimisation strategies and modern AI/ML add-ons, so any hands-on exposure to libraries such as PyTorch, NumPy, SciPy or TensorFlow will help you challenge them at the right level, though it is not mandatory. Availability is limited to two focus...

    $251 Average bid
    $251 Avg Bid
    15 bids

    ...narrative continuity before passing curated context into a citation aware LLM routing layer that prioritizes Gemini, OpenAI, then Anthropic, then Ollama local models, enforcing context bound generation and preventing hallucination outside retrieved evidence. Indexing is parallelized using ProcessPoolExecutor for efficient multi core utilization and automatically scales to distributed ingestion via PySpark when corpus size exceeds a configured threshold, enabling safe handling of 20k plus documents or 50GB class corpora, while the system is wrapped in a full MLOps backbone that integrates MLflow for experiment tracking of retrieval metrics, PPO reinforcement learning rewards, and parameter tuning, exposes Prometheus metrics for latency and retrieval monitoring compatible with Graf...

    $244 Average bid
    $244 Avg Bid
    14 bids

    Description: We’re looking for an experienced Data Engineer preferably based from Dubai to help build and manage data pipelines for a global platform. Most work is in Azure, using Azure Data Factory, ADLS, and Databricks. What you’ll do: Build and manage PySpark/Spark pipelines in Databricks Schedule and monitor pipelines in Azure Data Factory Optimize Databricks for better performance Keep code and documentation organized and clear Requirements: Experience with Azure cloud and Databricks Strong PySpark / Spark skills Experience building scalable, reliable data pipelines Details: Project-based, with potential to move to full-time Ideal for engineers who like building cloud-native pipelines

    $12 / hr Average bid
    $12 / hr Avg Bid
    20 bids

    ... • Read multiple flat-file formats (mainly CSV, with the occasional JSON). • Apply thorough data-cleansing rules—removing duplicates, enforcing data types, flagging out-of-range values, and normalising text fields. • Run validation checks so that only clean, schema-compliant rows proceed to the load step. I’m happy for you to choose the stack you are most efficient with—Python (pandas, PySpark), Talend, or another ETL tool—as long as the final solution is reproducible and can be triggered automatically (CLI, scheduled job, or cloud function). If you think aggregation or more advanced joins would improve the dataset, flag that as a future enhancement; for now, cleansing and validation are the must-haves. Deliverables 1. Well-docum...

    $26 Average bid
    $26 Avg Bid
    24 bids

    ...Azure Data Engineer to support and enhance our existing data platform on an ongoing basis. You should be strong in: Azure Data Factory (ADF) for building and maintaining ETL/ELT pipelines Azure Databricks and PySpark for large‑scale data processing Python for data engineering utilities, automation, and integration Delta Lakes/Lakehouse concepts, performance optimization, and troubleshooting Working with SQL‑based data sources, data warehousing, and BI integrations Responsibilities Design, build, and optimize data pipelines in Azure ADF and Databricks Develop and maintain PySpark and Python jobs for batch and near real‑time workloads Implement best practices for data quality, observability, and monitoring Collaborate with our internal team, follow existing standa...

    $10 / hr Average bid
    $10 / hr Avg Bid
    33 bids

    I am looking for an experience data engineer with 4-5 years of experience with Pyspark And Python handson experience. Experience with handling a complex data pipeline.

    $5 / hr Average bid
    $5 / hr Avg Bid
    6 bids

    ...Databricks Data Analyst and Data Engineer certifications and want a structured, hands-on tutoring program that also deepens my Snowflake skills. The goal is to become confident building end-to-end data pipelines, running analytics, and understanding platform architecture well enough to pass the exams and perform the work in practice. Focus areas Databricks • Data processing & analytics with PySpark/SQL and Delta Lake • Machine learning workflows inside the Databricks environment • Workspace, cluster, job, and Lakehouse architecture Snowflake • Core data-warehousing concepts and best practices • Query tuning and overall performance optimisation • Security features: RBAC, masking, encryption, and access policies How we can wor...

    $186 Average bid
    $186 Avg Bid
    48 bids

    ...Object Storage, Data Flow (Spark), and Data Catalog. * Solid understanding of Finance / Order-to-Cash (O2C) data entities and processes. * Knowledge of data modeling, lineage, and governance principles. * Familiarity with CI/CD and DevOps for automated deployments. Preferred Skills * OCI Data Integration certification. * Experience integrating Oracle Cloud ERP with OCI DI. * Knowledge of Python or PySpark for custom transformations. * Exposure to Data Science and ML pipelines leveraging OCI services. * Experience with monitoring tools like Grafana...

    $2069 Average bid
    $2069 Avg Bid
    5 bids

    ...guidance with embedding Genie via API into apps, Teams, or dashboards. • Train internal teams on Genie capabilities, administration, and operational readiness. Required Skills & Experience • Strong practical experience with Azure Databricks, Lakehouse architecture, Unity Catalog, SQL Warehouse. • Knowledge of Genie AI, foundational models, or Databricks conversational analytics. • Competency in PySpark, SQL, data modeling, and enterprise data engineering practices. • Familiarity with Azure ecosystem (Data Lake, Data Factory, DevOps). • Ability to translate business questions into NLQ-friendly dataset design. • Excellent communication and ability to work with cross functional data, BI, and business teams. Nice to Have • Experience with A...

    $2 / hr Average bid
    $2 / hr Avg Bid
    3 bids

    I have a Hadoop cluster holding several large data sets, and I need a seasoned PySpark developer who also writes rock-solid SQL. The immediate aim is to connect to the cluster (YARN/HDFS with Hive metastore), develop or refine PySpark jobs, optimise the accompanying SQL, and make sure everything runs smoothly end-to-end. You’ll receive access to a staging namespace plus a sample of the data. Once the logic checks out we’ll promote the code to the full environment. Deliverables • A clean, well-commented PySpark notebook or .py job that executes successfully on the cluster • The corresponding SQL script or view definitions ready for Hive or spark-sql • A concise README detailing execution steps, parameters, and expected outputs Accep...

    $75 Average bid
    $75 Avg Bid
    11 bids

    I need a reusable ETL framework built inside Databricks notebooks, version-controlled in Bitbucket and promoted automatically through a Bitbucket Pipeli...attached to any cluster. Acceptance criteria • Parameter-driven notebooks organised by layer. • Reusable GraphQL connector packaged as a .whl. • Bitbucket Pipelines yaml that runs unit tests, uses the Databricks CLI to deploy notebooks, and executes an integration test on commit. • Clear README detailing how to add a new API endpoint and where to place cleaning logic. Leverage native tools—PySpark, SQL, Delta Lake, dbutils—while keeping external libraries to a minimum and fully documented. Please share a brief outline of your approach and any relevant Databricks + Bitbucket CI experience s...

    $341 Average bid
    $341 Avg Bid
    114 bids

    I’m a beginner looking for a 1-on-1 Databricks instructor for a very hands-on, fast-paced 2-week program. Requirements: - Strong real-world Databricks experience - Hands-on Apache Spark (PySpark), SQL, Delta Lake - Real use case / mini project (end-to-end pipeline) - Live screen sharing, coding together - Beginner-friendly but practical (no theory-only) Goal: By the end of 2 weeks, I want to confidently build and understand a real Databricks data pipeline. Availability: 5–6 sessions per week, 1–1.5 hours per session Please share: - Your Databricks experience - How you would structure these 2 weeks - Your hourly rate Thanks!

    $20 / hr Average bid
    $20 / hr Avg Bid
    59 bids

    ...across multiple source systems. Build and optimize Foundry pipelines using Code Workbooks (PySpark, SQL, Scala) and Quiver. Support data integration, feature engineering, and pipeline debugging for production AI workloads. Implement security and permissions architecture aligned with enterprise governance. Help develop Foundry applications using Workshop, Contour, and Slate for analytics and decision-making. Guide on best practices for CI/CD, testing, and deployment within Foundry. Provide mentorship and troubleshooting support during live client engagements. Required Skills: Strong hands-on experience with Palantir Foundry (Ontology, Code Workbooks, Quiver, Workshop). Proficiency in Python, PySpark, and SQL. Experience with data modeling, transformation logic, and pipelin...

    $14 / hr Average bid
    $14 / hr Avg Bid
    20 bids

    Need a strong streaming experience person to develop design deploy Pyspark publishing and upserting job in EMR with Spark, MongoDB(documentDb) connector, AWS EMR step functions, Cloud watch, docker, Kafka cluster architecture, Airflow dags, Gitlab, Pycharm, Cursor AI IDE etc needed for environment experience

    $11 / hr Average bid
    $11 / hr Avg Bid
    39 bids

    ...patterns that Databricks loves to test. • Fresh practice questions (or a curated question bank) with detailed explanations so I understand not just the right answer but the thinking process. • At least one full-length mock exam under timed conditions followed by a debrief on weak areas and strategies to avoid common pitfalls. I work mainly in the Databricks notebook environment with Python, PySpark, and SQL, so please weave real-world examples into the prep. I’m flexible on session times and frequency; we can agree milestones and refine the plan as we go. If you’ve already helped others pass this exam—or you hold the certification yourself—tell me how you’d tackle my study roadmap and what materials you’d bring to the table. I...

    $46 Average bid
    $46 Avg Bid
    2 bids

    ...actual medicines and would map once the inconsistencies are ironed out, so I want the process to be fully automated, driven by a robust auto-correct algorithm rather than manual review. Remaining 0.1% could be non medical entries, and need to be deleted. I am open to proven techniques—fuzzy matching, phonetic hashing, Levenshtein, word embeddings, or a hybrid—as long as they scale. Python, pandas, PySpark, or any other big-data friendly stack is fine, provided the final solution is reproducible and well documented. Deliverables • Clean, executable scripts (Jupyter notebook or .py) that ingest both files, normalise product names, detect duplicates, and output a one-to-one mapping table. • A brief README explaining dependencies, algorithm logic, and how ...

    $607 Average bid
    $607 Avg Bid
    39 bids

    ...Infrastructure Microsoft Azure (Functions, Logic Apps, Service Bus, Blob Storage, Data Factory, Azure DevOps) AWS Cloud Docker, Kubernetes RabbitMQ CRM, ERP & Enterprise Platforms Microsoft Dynamics CRM 365 Dynamics Business Central Sage CRM NopCommerce Sitefinity v12.2 Umbraco v8.0 DotNetNuke v4.0 Python, AI & Advanced Solutions Python, Django, Flask, Pyramid REST APIs, WebSockets PySpark AI Email & Chatbot Solutions Data Science & Analytics CMS, E-Commerce & Web Platforms WordPress, Joomla, Drupal Prestashop PHP-based systems BI, Finance & Business Support Power BI Advanced Excel Accounting, Finance & Bookkeeping Data Entry & Business Reporting MS Office Suite Tools & Delivery Methodology Git (Version Control) N...

    $5 / hr Average bid
    $5 / hr Avg Bid
    20 bids