
Closed
Posted
I have a collection of unstructured data—mixed text documents and images—and I need an unsupervised learning workflow that reliably flags unusual or suspicious examples. The aim is purely anomaly detection; no labels are available and none can be added, so clustering or classification won’t help here. Here’s the flow I have in mind: • Data preparation: consistent preprocessing for both modalities (tokenisation or embeddings for text, feature extraction for images). • Model development: an unsupervised architecture such as auto-encoder, variational auto-encoder, deep clustering, or another approach you can justify for anomaly detection. • Evaluation: quantitative metrics (reconstruction error distributions, AUC, or similar) plus a concise report explaining thresholds and decision logic. • Deliverables: clean, well-commented Python code (ideally PyTorch or TensorFlow/Keras), reproducible environment files, and a short README so I can retrain or fine-tune later. I will supply a representative sample to start; please keep the design modular so it scales once the full dataset arrives. Let me know any additional dependencies you anticipate, along with an outline timeline for data exploration, model iteration, and final validation.
Project ID: 40665435
42 proposals
Remote project
Active 1 hour ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
42 freelancers are bidding on average $18 USD/hour for this job

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Python, and similar tools. I have worked with pytorch, and tensorflow to develop DL models, .I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
$20 USD in 40 days
7.6
7.6

With my vast experience as an AI and Cloud Developer, specializing in building intelligent applications, I am confident I can execute your unsupervised anomaly detection project with excellence. I have a strong background in Algorithm, Deep Learning, and Machine Learning which are key for developing models like auto-encoders, variational auto-encoders, and deep clustering system that we will be needing for this project. One thing that sets me apart is my holistic approach to application development. Beyond just creating AI models, I also excel at building backend systems, web dashboards and APIs - all of which will come in handy for a project of this scale and scope. Furthermore, I'll leverage cloud infrastructure to ensure your project scales smoothly as you add the full dataset. In addition to this technical prowess, I am also known for delivering clean and well-documented code along with comprehensive reports that help my clients understand and take ownership of the platforms once the project has been completed. With me on board, you can expect not just an efficient workflow but also clear communication throughout the process to ensure all details are understood and deadlines met. Should any additional requirements arise from your end, rest assured I can handle them too.
$25 USD in 40 days
6.5
6.5

I am an expert statistician, Research Writer, and data analyst with more than eight years of experience. I have full command of Excel analysis, SPSS, STATA, R LANGUAGE, AND PYTHON. I am an expert in creating time series prediction models, working with survey data, conducting marketing analysis, building estimators, and medical analysis. I am a perfect match for your project share other details of the work so I can start working on your project. Will complete task on time.
$15 USD in 10 days
5.6
5.6

Hi,I am a seasoned Applied Machine Learning Engineer(6+ yoe)& experienced in production unsupervised anomaly-detection pipelines,multimodal embeddings,similarity modelling & sparse/no-label environments -Built an industrial predictive-maintenance anomaly engine for real plant vibration streams where labelled failures were unavailable: removed operating-context effects,generated rolling/EWMA residual features,Mahalanobis health indicators & unsupervised deviation scores to separate genuine degradation from normal load-driven vibration -Also engineered a production multimodal image embedding pipeline using IResNet/AdaFace representations,pgvector similarity search & unsupervised identity clustering,solving niche issues such as cluster drift,threshold calibration,duplicate observations & conflicting cluster assignments -For your text+image data,I propose pretrained semantic/image embeddings -> modality-specific normalization -> Deep SVDD/VAE or robust latent-density scoring -> calibrated anomaly fusion -A key technical issue is that reconstruction error is often poorly comparable across modalities;I’ll use per-modality robust score calibration (median/MAD or empirical tail probabilities) before fusion so image variance cannot dominate textual anomalies -Evaluation will include score distributions,contamination sensitivity,stability tests,synthetic holdout anomalies where appropriate,threshold rationale & explainable nearest-normal comparisons
$15 USD in 40 days
4.3
4.3

Hello, Your requirement for unsupervised anomaly detection on mixed text and image data without any labels shows you understand that traditional classification approaches cannot apply here. Choosing the right autoencoder architecture and reconstruction error thresholds is what makes this workflow actually flag suspicious examples reliably. I'd start by exploring your representative sample to determine whether separate modality-specific encoders or a joint embedding space better captures anomalies in your data distribution. I have debugged and tuned deep learning models in Keras and built NLP systems including text classification APIs, so I know how to design variational autoencoders and evaluate them using reconstruction error distributions rather than supervised metrics. I will deliver modular PyTorch or TensorFlow code with reproducible environment files, clear threshold documentation, and a README enabling future retraining. The architecture will scale cleanly when your full dataset arrives. Do your text documents and images share contextual relationships, or should each modality be processed independently for anomaly scoring? Share the sample data and I will outline the preprocessing and model selection strategy before development begins. Adrian
$15 USD in 40 days
2.9
2.9

Hello, I can develop a modular and reproducible anomaly-detection workflow for your mixed text and image data using Python and PyTorch. My background includes Python, machine learning, AI-powered data pipelines, feature engineering, and large-scale data processing. I focus on building practical, maintainable systems rather than delivering only an experimental notebook. I also have a 5-star Freelancer rating with 100% on-time and on-budget delivery. For this project, I will: Inspect the representative sample and assess data quality and structure. Build consistent preprocessing pipelines for both text and images. Generate semantic text embeddings using an appropriate pretrained transformer. Extract image features using a suitable pretrained vision model. Develop separate modality-specific anomaly scores before testing a combined scoring method. Compare autoencoders, variational autoencoders, and suitable classical anomaly-detection baselines. Calibrate transparent thresholds from reconstruction-error and anomaly-score distributions. Return ranked suspicious examples with text, image, and combined anomaly scores. Keep preprocessing, training, thresholding, and inference modular so the workflow can scale when the complete dataset arrives. Since no ground-truth labels are available, conventional AUC cannot be calculated reliably. I will therefore evaluate the workflow using reconstruction-error distributions, score stability, synthetic anomaly injection, sensitivity analysis, and inspection of the highest-ranked anomalies. The final report will clearly explain the selected threshold and decision logic without presenting misleading accuracy metrics. You will receive: Clean, well-commented Python code Separate training and inference scripts Saved model weights or checkpoints Reproducible environment and dependency files Configuration-based model and threshold settings An evaluation report with quantitative results and visualisations A concise README covering setup, execution, retraining, and fine-tuning Proposed timeline: Days 1–3: Data exploration and preprocessing Days 4–7: Feature extraction and baseline anomaly models Days 8–10: Autoencoder/VAE development and model comparison Days 11–12: Threshold calibration and scalability testing Days 13–14: Documentation, revisions, and final handover My bid is $18 per hour, with an estimated commitment of 40–45 hours over 14 days. I will provide regular progress updates and deliver components incrementally so you can review the preprocessing, model results, and anomaly-scoring logic before final handover. Before starting, I would confirm the approximate dataset size, file formats, whether each text document corresponds to an image, the available training hardware, and whether pretrained model downloads are permitted. Best regards, Robert AI Data & Automation Specialist | Python
$18 USD in 25 days
2.7
2.7

Hi, I can build a modular unsupervised anomaly-detection pipeline for your mixed text and image data using Python and PyTorch. I’ll cover preprocessing, embeddings/feature extraction, autoencoder or VAE-based modeling, anomaly scoring, threshold selection, evaluation, and visualization. The final package will include clean code, requirements/environment files, README, and a short report. I can start with your sample dataset, benchmark a few suitable approaches, and select the most reliable one before scaling. Regards, Abdul Samad
$15 USD in 40 days
1.8
1.8

⚠️ IF YOU'RE NOT HAPPY YOU DON’T PAY ⚠️ I think we’re a strong fit for your project. I specialize in Python-based ML pipelines for unstructured text and image data, with modular preprocessing and reproducible training workflows. For this anomaly-detection system, I’d encode text and images separately, then evaluate an unsupervised approach such as autoencoders/VAEs or embedding-distance methods based on the sample distribution rather than forcing one model upfront. I’d keep preprocessing, feature extraction, scoring, thresholding, and evaluation as separate components so the pipeline scales cleanly when the full dataset arrives. The main technical issue is evaluation without labels, so I’d define threshold logic using score distributions, synthetic perturbation tests, and stability analysis. Multiple 5-star reviews on machine learning, Python, and data-analysis projects. I’d love to chat about your project! The worst that can happen is you walk away with a free consultation. Regards, Chris
$20 USD in 40 days
1.6
1.6

As a seasoned AI and ML expert with over two decades of hands-on experience, I am confident in my ability to develop an exemplary unsupervised anomaly detection model for your unstructured data. From my previous roles as a Chief Technology Officer and software engineer, I have a diverse range of skills that include artificial intelligence, machine learning, data science, and predictive analytics—skills that are crucial to the success of this project. My understanding of technologies like PyTorch and TensorFlow/Keras, combined with the proficiency I have acquired as part of my full stack web and mobile app development experience distinguish me as an ideal professional for the task. I can ensure that all codes will be clean, well-commented following a modular design, and will yield a reproducible environment. Moreover, my comprehensive grasp on overall system optimization and cloud deployment equips me not only to deliver the anomaly detection model you require but to also provide valuable advice and strategies for its future retraining or fine-tuning. Hiring me would give you the advantage of a technical consultant who specializes in AI end-to-end- solutions. Let's do this together!
$15 USD in 40 days
0.0
0.0

The data preparation for mixed text and images is the part that usually breaks on jobs like this, so I will use Sentence-BERT embeddings for the text and extract features from images using a pre-trained ResNet50, then combine these into a single feature vector per data point. The unsupervised learning workflow will use a Variational Auto-Encoder (VAE) as it is good for learning compact representations and identifying outliers based on reconstruction error, so it fits anomaly detection directly without needing labels. I will not use deep clustering. Deep clustering requires inferring cluster assignments, which introduces an extra layer of abstraction not needed for pure anomaly detection when the goal is simply to find what deviates from the norm. For evaluation, I will look at the distribution of reconstruction errors and use the interquartile range to set a dynamic threshold, reporting the AUC on a small, held-out set of known anomalies if you can provide it. I am a Preferred Freelancer here, and I have not missed a deadline or gone over an agreed price yet. What is the maximum size of the data collection you expect to process? Once I have that, I will send a revised timeline and confirmation of the VAE architecture.
$25 USD in 7 days
0.0
0.0

Hi, I’d be happy to build your unsupervised anomaly detection pipeline for both text and image data. I have experience with Python, machine learning, deep learning, and models such as autoencoders, and I can design the workflow to remain modular and scalable as your full dataset grows. My approach would include: * Data exploration and consistent preprocessing for text and images * Feature extraction/embeddings suitable for each modality * Autoencoder or VAE-based anomaly detection, with the architecture selected based on the sample data * Reconstruction-error analysis and statistically justified anomaly thresholds * Quantitative evaluation using suitable metrics and visualizations * Clean, well-commented PyTorch/TensorFlow code * requirements/environment file and concise README for reproducibility I’ll structure the project so preprocessing, training, inference, and thresholding are modular and easy to modify later. I can start with your representative sample, establish a strong baseline, iterate on the model, and then perform final validation. I’d be glad to review the sample data and propose the most suitable architecture before implementation.
$16 USD in 40 days
0.0
0.0

As a well-rounded software engineer experienced in delivering scalable applications, I'm confident that I have the interdisciplinary skills and adaptability to build an effective unsupervised anomaly detection model for your unique project. My extensive knowledge in multiple languages, including Python (which has powerful ML libraries like PyTorch and TensorFlow/Keras) ensures that the final solution will be well-commented, clean, and maintainable. One of my core strengths lies in translating complex business goals into efficient software solutions. For example, I've recently built a Real-Time Automated Supply Chain System using MQTT, WebSocket, AWS, and other related technologies. The insights I gained from handling a large volume of unstructured data in real-time will be highly transferable to your use case. Moreover, apart from just delivering the desired code and environment files, I’m committed to providing you with a concise report that comprehensively explains the thresholds and decision logic specific to your data. This way, even after the engagement terminates, you'll have a clear understanding of how to retrain or fine-tune the model as your dataset expands or changes. With me on board, you can expect not just a one-off solution but also an evolved system that can scale alongside your specific needs.
$15 USD in 40 days
0.0
0.0

Hi, I understand your requirement for an unsupervised anomaly detection system on mixed text and image data without labeled examples. I have recently studied Machine Learning and Deep Learning thoroughly during my 6th semester and achieved 86/100 (A Grade). I have practical understanding of unsupervised learning, neural networks, feature extraction, and model evaluation. For your workflow, I can help with: • Preprocessing and preparing both text and image modalities • Extracting meaningful embeddings/features from documents and images • Building an unsupervised anomaly detection pipeline using approaches like Autoencoders, Variational Autoencoders, or deep feature-based methods • Setting proper anomaly thresholds using reconstruction error and evaluation metrics • Providing clean Python code (PyTorch/TensorFlow), documentation, and reproducible setup I understand that since no labels are available, the main challenge is designing a model that can learn normal patterns and reliably identify unusual samples. I will focus on creating a modular solution that can scale when your full dataset is provided. I would be happy to review your sample data and suggest the most suitable approach. Thank you.
$15 USD in 35 days
0.0
0.0

Hi, I am a software engineer with over 16 years of experience building Python-based data and machine-learning systems. I can develop a modular anomaly-detection workflow for your mixed text and image dataset, from reproducible preprocessing through scoring, threshold selection, and validation. I would first explore the sample and establish strong baselines using pretrained text and vision embeddings with modality-specific detectors. I would then compare these with an autoencoder or VAE where the dataset size supports meaningful training, calibrate the anomaly scores, and combine them into an explainable final ranking. Since no labels exist, evaluation will use score distributions, stability checks, held-out reconstruction results, and controlled synthetic anomalies; AUC can be included only if a defensible proxy evaluation set is available. You will receive clean PyTorch/Python code, pinned environment files, a README, and a concise report covering dependencies, thresholds, and decision logic. I anticipate roughly 2–3 days for exploration, 4–6 days for model iteration, and 2–3 days for validation and documentation. What are the approximate dataset size and dominant document/image formats? I would be glad to discuss the details and start with your representative sample.
$25 USD in 20 days
0.0
0.0

Hi, this is a great fit for my background in unsupervised ML and multimodal data. Approach: I'd avoid a single joint text+image model (hard to debug with no labels to validate against). Instead — separate embeddings per modality (sentence-transformers for text, pretrained CNN/ViT features for images), each scored independently for anomalies, then combined at the score level. This keeps the pipeline modular and lets either branch be retrained on its own as your dataset grows — matching your scalability requirement directly. I'd start with a fast baseline (Isolation Forest / PyOD) on the embeddings to sanity-check things quickly, then move to an autoencoder or VAE as the production model, justified against your actual sample rather than picked blind. Evaluation: since there's no ground truth, I'll document a statistically grounded threshold (percentile/EVT-based) on the anomaly-score distribution, run stability checks across re-runs, and inject synthetic anomalies to confirm the pipeline catches obvious outliers before trusting it on subtle ones — all summarized in a short report. Deliverables: clean, modular PyTorch code, environment file, README for retraining. Timeline: once I have the sample — exploration + baseline (~6-8h), main model + iteration (~12-15h), evaluation/report (~5-6h). Happy to work within your $15/hr and track hours transparently. Ready to start as soon as the sample is shared — happy to answer any questions first.
$15 USD in 40 days
0.0
0.0

Hello, Your project closely matches my experience in Python, deep learning, computer vision, image processing, and experimental machine learning. I have hands-on AI research experience, including an AI/IoT project for predicting mechanical ventilator failures and research on brain MRI classification using Vision Transformers. This work involved preprocessing experiments, model evaluation, analysis of model behavior, and interpretability. For this project, I would first analyze the representative sample rather than assume a specific architecture. Since the data contains text and images without labels, I would establish suitable representations for each modality and compare appropriate unsupervised approaches such as autoencoders, VAEs, and embedding-based methods. The workflow would cover data preprocessing, image feature extraction, text embeddings, anomaly scoring, threshold selection, multimodal score integration, and quantitative/visual validation. I would ensure the evaluation strategy is appropriate for an unlabeled setting and deliver clean, modular PyTorch code, a reproducible environment, and clear documentation covering the methodology and decision logic. I would be happy to review the sample and establish a reliable baseline first. Best regards, Arwa
$18 USD in 20 days
0.0
0.0

Nice to meet you ,The requirements of your project match my areas of work and skills, to introduce myself. My name is Anthony Muñoz and i am the lead engineer for DS Pro IT agency. I have worked for over 10 years as a Full-Stack and software development engineer and have successfully done multiple jobs. It will be a pleasure to work together to make your project. Feel free to discuss about the project with me, greetings.
$18 USD in 40 days
3.6
3.6

I have hands-on experience with Python, Machine Learning, Deep Learning, and Computer Vision, and I’m comfortable working with unlabelled datasets and building end-to-end ML pipelines. For this project, I can work with the text and image modalities separately, handle the required preprocessing/feature extraction, and develop an unsupervised anomaly detection approach such as an autoencoder or variational autoencoder. I can also evaluate reconstruction errors and define practical anomaly thresholds based on the data. I focus on clean, modular, and reproducible Python code, with proper documentation so the workflow can be extended when the complete dataset is available. I’d be happy to start with the representative sample and iterate based on the results.
$16 USD in 25 days
0.0
0.0

Unsupervised anomaly detection often founders on defining "unusual" without labels, especially across mixed data types. I build custom AI agents on the Claude Agent SDK with autonomous perceive-plan-act-reflect loops, and I create custom RAG/retrieval pipelines. I also engineer multi-agent orchestration systems. For your mixed text and image data, I would design a custom multi-modal embedding approach, then build an anomaly detection system based on reconstruction error or density estimation. I see you're planning for modularity and scalability for a full dataset. I run self-hosted infra with Postgres, cron, and queueing, and I build Playwright scrapers for data collection, so I understand robust system design. My rate is $15 per hour, capped at 20 hours a week. I would start with a short scoped block to build the data ingestion pipeline for your sample and test initial embedding strategies for both modalities, so you see output before committing further hours. Please share your representative sample. I will analyze its structure and propose a specific architectural approach for the autoencoder or VAE component.
$15 USD in 20 days
0.0
0.0

Hello, I’m an Informatics professional with a Master’s degree and hands-on experience in Python, machine learning, data preprocessing, model development, and evaluation. Your project is a strong match for my background. I can develop a modular unsupervised anomaly detection workflow for both text and image data, including preprocessing, feature extraction, autoencoder/VAE-based modeling, anomaly scoring, threshold analysis, and quantitative evaluation. I can provide clean and well-documented Python code, a reproducible environment, README documentation, and a concise report explaining the methodology, thresholds, and decision logic. I’m also comfortable working iteratively, starting with your representative sample before scaling the pipeline to the full dataset. I would be happy to review the sample data and propose the most suitable architecture and implementation approach. Best regards, Roymond
$15 USD in 40 days
0.0
0.0

Kiambu, Kenya
Member since Aug 18, 2026
$250-750 USD
$15-25 USD / hour
$250-750 USD
$250-750 USD
€30-250 EUR
₹12500-37500 INR
$250-750 USD
$8-15 USD / hour
$25-50 USD / hour
$15-25 USD / hour
₹1500-12500 INR
₹75000-150000 INR
$15-25 USD / hour
$7-10 USD / hour
$2000-6000 HKD
$100-150 USD / hour
$10-30 CAD
$10-30 USD
₹12500-37500 INR
£50000-100000 GBP