
Closed
Posted
Rubric Task Author — Project Bulls Eye (Contract, Paid Per Task) About the Role AI models are already being used for real professional work — writing reports, answering questions, building analyses. But they're not reliable enough yet: they still get things wrong in ways that matter, and right now nobody can consistently find out where. We're looking for domain experts to help find those breaking points. You'll build realistic, professional-grade deliverables in your area of expertise, then build the rubric that grades them — creating prompts hard enough that today's best AI models still get wrong, but realistic enough that a working professional would genuinely need them answered correctly. This work spans 67 occupations across 16 domain groups, including Healthcare, Finance & Accounting, Legal & Compliance, Engineering, Supply Chain & Procurement, and Business Operations. What You'll Do Design a realistic professional task from your field — the kind of question, report, or analysis a working professional would actually be asked to produce. Write the prompt itself: the realistic, hard question or scenario a professional would need answered. Build the deliverable yourself, to the standard a qualified professional would expect. Build a rubric that precisely and objectively grades that deliverable — capturing exactly what separates a correct, professional-quality answer from a flawed one. Stress-test the task against current AI models to confirm it's genuinely hard — the model should get it wrong or fall short in a meaningful way. Pass the project assessment required to qualify for task work. Submit your task for review and validation. What We're Looking For Real professional experience in one or more of the 16 domain groups (Healthcare, Finance & Accounting, Legal & Compliance, Engineering, Supply Chain & Procurement, Business Ops, or similar). Deep enough expertise to know what a genuinely hard, realistic professional scenario looks like in your field — not a textbook question, but something with the ambiguity, edge cases, or judgment calls real work involves. Ability to write clearly and define objective, defensible grading criteria — you're not just producing an answer, you're specifying what "correct" means. Comfort working independently through a structured, guided workflow. Requirements to Get Started Complete and pass the Project Bulls Eye project assessment to qualify for task work. For each task, write an original prompt and build out the full deliverable + rubric as described above. Pay Structure $30 per task, paid per task rather than by the hour. Payment is issued only after your task is validated and passes review — meaning it clears the required quality checks and is confirmed to meet the project's standards before it's accepted. Tasks that don't pass validation are not eligible for payment; you're welcome to revise and resubmit where applicable. Engagement Details Contract / freelance, task-based. Work through the guided workflow in the Project Bulls Eye learning hub, then pass the project assessment before submitting your first task. Best suited for professionals who want flexible, self-directed contract work tied to their existing domain expertise. To apply, tell us which of the 16 domain groups and occupations best match your background, and briefly describe a real, hard professional scenario from your field that you think would stump an AI model today.
Project ID: 40674193
19 proposals
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
19 freelancers are bidding on average $7 USD/hour for this job

Hello, {{{ I HAVE 10+ YEARS OF PROFESSIONAL EXPERIENCE IN SOFTWARE ENGINEERING, AI, FULL-STACK DEVELOPMENT, API INTEGRATION & SYSTEM ARCHITECTURE AND I CAN SHOW YOU RELEVANT WORK }}} I have carefully reviewed the Project Bulls Eye requirements and understand that the focus is on creating realistic, difficult professional tasks with objective, defensible rubrics. My strongest match is Engineering / Technology, AI & Software Development. A realistic challenging scenario from my field would be asking an AI model to diagnose a production system with intermittent API failures, queue backlogs, database contention and incomplete logs, then identify the root cause, prioritize remediation and provide a technically justified recovery plan. This tests reasoning, trade-offs and practical engineering judgment rather than simple coding knowledge. I am comfortable creating the complete professional deliverable, detailed grading rubric and stress-testing the task against AI models. I can work independently within your structured workflow and revise tasks based on validation feedback. I eagerly await your positive response. Thanks, Christina
$10 USD in 30 days
4.9
4.9

I have experience crafting clear, effective rubrics that improve grading consistency and support AI training data quality. Your project’s focus on precision and clarity in task design aligns well with my skills. I have developed educational rubrics used in AI-enhanced learning platforms, ensuring tasks are measurable and outcomes straightforward to assess. My approach involves closely aligning rubric criteria with learning objectives and iteratively refining tasks based on feedback to ensure clarity and reliability. Do you have specific guidelines or examples for the rubric tasks you want created?
$5 USD in 7 days
0.0
0.0

Hi, I’ve reviewed your project Expert Rubric Task and understand what you’re looking to achieve. I can help you build and deliver the complete solution with a focus on clean implementation, performance, reliability, and a polished user experience. I’ll first understand your existing setup and requirements, then suggest the most practical approach without adding unnecessary complexity. I’m available to start quickly and can work with your preferred technologies and workflow. Send me a message with the details, and let’s discuss the project, timeline, and next steps.
$5 USD in 40 days
0.0
0.0

Hi , You need an expert in Research Writing, Prompt Engineering, AI Content Creation, AI Automation, AI Research, AI Model Development and AI Training Data, and I have a tailor-made solution ready for you. Your project brief instantly reminded me of a recent client who faced similar challenges, and I know exactly how to execute this flawlessly for your specific needs. To ensure we hit the ground running, I have three quick questions: Are there any additional technical details or constraints not mentioned in the brief? What is the primary hurdle currently blocking your progress on this? What is your strict timeline for completion? Why trust me with your project? The Record: 250+ Projects. 6+ Years. 100+ consecutive 5-star reviews. The Standard: Zero misses. I don’t just finish the job; I guarantee flawless execution. The Availability: Full-time freelancer, online 9 AM - 9 PM EST. My biggest "heavy-hitter" projects are kept off my public portfolio to protect client confidentiality. Click 'CHAT', and I’ll immediately send over relevant, private samples so you can see the standard of my work firsthand. Best regards, Muhammad Arsalan
$10 USD in 28 days
0.0
0.0

Drawing from my strengths as a software engineer specializing in cloud technology, I believe I have the unique skill set to approach this task authoring project with a fresh, informed perspective. The rubric-building part is an area in which I excel, and I can offer you cutting-edge, efficient and reliable solutions that would undoubtedly enhance user experiences. Additionally, my experience in researching and writing goes hand-in-hand with crafting prompts that are not only realistic but also challenging—an essential component for stumping today's AI models. Working through the Project Bulls Eye guided workflow independently is comfortable for me as I've been solving complex problems autonomously throughout my career. This attribute is crucial as it reflects on my ability to submit top-notch tasks that genuinely stress-test current AI models. Overall, not only am I prepared for flexible, self-directed work to be executed at the highest standards, but my unique combination of expertise also positions me as an incredible fit for your project. With me on board, rest assured you'll have deliverables and rubrics that experts in various occupancies across industries can relate to closely while effectively separating correct answers from flawed ones. Thank you for considering my application.
$5 USD in 40 days
0.0
0.0

Your focus on creating realistic professional tasks that expose where AI models fail is particularly interesting because the strongest evaluations aren't textbook questions—they require judgment, ambiguity, edge-case handling, and a defensible definition of what “correct” actually means. I can contribute by designing professional-grade scenarios, producing the expected deliverable, and building objective rubrics that distinguish a genuinely strong answer from one that merely sounds convincing. I’m comfortable breaking complex problems into measurable criteria, identifying critical omissions, and stress-testing whether a task actually requires meaningful reasoning. The areas I can contribute most strongly to include AI/technology, business operations, automation, software development, and AI-powered workflows. A strong example would be an AI automation scenario where the model must evaluate conflicting business requirements, API limitations, failure states, data-handling considerations, and propose a reliable implementation rather than simply generating generic workflow steps. I’m also comfortable working independently through a structured assessment and revision process, with quality and validation taking priority over task volume. Expert question: would you prefer contributors to focus on one primary occupation/domain throughout the engagement, or can we submit tasks across multiple closely related technical and business areas? George
$7 USD in 40 days
0.0
0.0

Hello, Warm greetings from Syndell! We understand you're building rubric tasks to identify where AI models fail on real professional work - creating scenarios realistic enough that working professionals need them solved correctly. We bring prompt engineering expertise and deep experience in AI evaluation frameworks, allowing us to craft tasks with the precision this work demands: questions that are both authentically professional and strategically difficult. Which domain or occupational group is your priority for these initial three tasks? We are a full-service software, AI, and digital-marketing agency with 5-star reviews on Clutch and GoodFirms, counted among the top-rated companies for our services on Freelancer.com. We would love the chance to discuss your vision in more depth and demonstrate, with examples from our portfolio, why we are the right fit. Hoping to speak soon. Thanks! PS: The final time and cost may be subject to change based on our discussion.
$2 USD in 7 days
0.0
0.0

Hi, I am a professional English <> Arabic translator with experience in medical and technical documents. I understand the Casgevy topic requires accurate research and clear, professional writing for pharmacy professionals. I have strong research skills and I can deliver 2-3 informative pages that are well-structured, factual, and easy to read. What I offer: - Accurate translation and writing with medical terminology - Fast delivery within 7 days - 100% original, well-researched content I would love to help you with this project. Let's discuss the details. Best regards, [Your Name]
$5 USD in 40 days
0.0
0.0

I am a Biology Lecturer and experienced educator with a strong background in designing challenging, concept-driven biology tasks and evaluating whether answers are scientifically accurate, logically sound, and professionally appropriate. My strongest fit is the Healthcare / Life Sciences domain, particularly Biology education and scientific reasoning. I can create realistic professional scenarios that require more than textbook recall, including situations involving experimental interpretation, physiology, genetics, data analysis, conflicting evidence, and scientific reasoning. For example, an AI model may correctly describe a biological mechanism in general but fail when required to interpret an unfamiliar experimental dataset, distinguish correlation from causation, identify a scientifically meaningful control, and justify its conclusion from the available evidence. I can turn such scenarios into rigorous prompts, complete professional deliverables, and objective rubrics that clearly distinguish a genuinely correct answer from a plausible but scientifically flawed one. I am comfortable working independently, following structured evaluation criteria, stress-testing AI outputs, and refining tasks based on review feedback. I would be glad to complete the required assessment and contribute reliable, high-quality evaluation tasks on an ongoing basis.
$8 USD in 20 days
0.0
0.0

Hi, I'm very interested in the Rubric Task Author position for Project Bulls Eye. I worked as Information Secretary for PTI for 2.5 years where I used ChatGPT, Gemini and Claude daily for content creation, research and communication. I have hands-on experience with AI tools and prompt engineering. I can create realistic professional tasks and build rubrics to identify where AI models fail. I am detail-oriented and can deliver high quality work on time. I can start immediately. Thank you for considering my proposal.
$8 USD in 40 days
0.0
0.0

Hi, My background is in bioinformatics and genomics research, and I’m currently working on genome assembly, genome annotation, transposable element analysis, ATAC-seq, RNA-seq and evolutionary/selection analyses. The domain group that best matches my background would be Healthcare/Life Sciences and Data Analysis/Research, depending on the available occupation categories. One type of task I think could genuinely stump an AI model is a real genome-analysis problem where the input files contain inconsistent annotations or unexpected formats. I’m comfortable creating the task, solving it myself to establish the correct answer, and then building a detailed rubric that objectively checks whether the AI got the important biological and technical details right. I’m interested in the paid task-based work and would be happy to complete the Project Bulls Eye assessment. Thank you.
$5 USD in 25 days
0.0
0.0

Hi, I’ve reviewed your project and understand that you need help with creating realistic professional tasks and rubrics for AI assessment. I can assist you with designing intricate scenarios in the Finance & Accounting domain, ensuring they reflect genuine work challenges that would stump AI models. My experience in financial analysis and report writing will allow me to craft prompts that not only test AI but also mirror real-world complexities. Additionally, I can build precise grading rubrics that define what constitutes a correct response, capturing the nuances that differentiate professional-quality work from flawed outputs. I will maintain clear communication throughout the process to ensure the final result meets your expectations. I'd love to chat about your project! The worst that can happen is you walk away with a free consultation. Regards, JaniceR92
$4 USD in 7 days
0.0
0.0

Hi, I have experience with Research Writing and Prompt Engineering. I can create clear, detailed rubrics and expert-level AI tasks that are easy to follow and grade. I understand how to write tasks that test critical thinking and ensure quality results. I am reliable, deliver on time, and follow instructions carefully. I would love to help with this project. Let's discuss the details. Thank you!
$5 USD in 40 days
0.0
0.0

My strongest domain areas are Engineering, Supply Chain & Procurement, Business Operations, and AI/optimization-related problems. A realistic difficult task I could design is a resource-allocation and scheduling scenario where limited resources must be assigned across competing jobs while satisfying capacity, timing, priority, service-level, and cost constraints. An AI model may produce a convincing answer while violating one or more coupled constraints, making it suitable for objective rubric-based evaluation. I can create the original professional prompt, reference deliverable, precise grading rubric, and stress-test the task against current AI models. I understand the compensation is $30 per validated task and that I must first complete the Project Bulls Eye assessment. I am ready to start immediately. Best, Mazyar
$5 USD in 10 days
0.0
0.0

Ethan here, from South Africa. Your project immediately caught my eye. I'm really excited to partner with you. With my extensive background in Finance & Accounting, I can design a realistic task that challenges AI models. For instance, I once crafted a complex financial analysis scenario that required deep judgment on asset valuation, which stumped even seasoned professionals. To tackle this project, I would create nuanced prompts that reflect real-world complexities, ensuring they’re challenging yet relevant. I’d focus on building clear grading rubrics that captivate what constitutes a professional-quality response, allowing for objective evaluation. I am committed to delivering high-quality work by rigorously testing each task against AI capabilities to confirm its difficulty. What you truly need is not just accurate tasks but those that reveal where AI struggles in professional scenarios. Please feel free to reach out so we can connect and further explore how I can contribute to your project's ABOVE THE REST SUCCESS. Kind regards, Ethan
$3 USD in 8 days
0.0
0.0

My strongest fit is Business Operations and Engineering/Software: AI workflow automation, incident diagnosis and acceptance testing. I understand payment is $30 per validated task, not the displayed hourly rate, and will complete the assessment and revisions. A realistic hard scenario from my field: an automation dashboard reports success, yet downstream customer or billing records are missing. Several layers show green because each validates an HTTP 200, a queued job or a stale cache rather than final persistence. The task would require reconciling live process state, configuration, API logs, retry/idempotency behaviour and database evidence; identifying the exact failure boundary; separating fact from plausible narrative; and proposing a safe fix with a falsifiable acceptance test. The deliverable would be a root-cause report and recovery plan. Its rubric would reward evidence-linked claims, separation of upstream acknowledgement from downstream completion, duplicate-event handling, safe rollback and measurable verification. It would reject unsupported certainty, “reset everything” fixes and tests that repeat the same misleading success signal. I can create original prompts, reference deliverables and objective grading criteria, stress-test them against current models, document meaningful failure modes and iterate after review. I would start with Business Operations / software-operations tasks and avoid domains where I lack real professional expertise.
$8 USD in 10 days
0.0
0.0

atlanta, United States
Payment method verified
Member since Oct 24, 2019
$2-8 USD / hour
$10-30 USD
$2-8 USD / hour
$10-30 USD
$2-8 USD / hour
$15-80 USD / hour
₹600-1500 INR
₹75000-150000 INR
$250-300 USD
₹12500-37500 INR
$250-750 USD
₹1500-12500 INR
$100-250 USD
$2-8 USD / hour
$15-25 USD / hour
₹1500-12500 INR
₹1500-12500 INR
$20000-50000 USD
₹12500-37500 INR
$2-8 USD / hour
₹500000-1000000 INR
₹12500-37500 INR
min £36 GBP / hour
₹12500-37500 INR
$30-250 USD