
Open
Posted
•
Ends in 3 days
Paid on delivery
I have a collection of patient medical records already cleaned and stored as CSV files, and I’m ready to turn them into an accurate diabetes-risk prediction tool. My preference is to work with a Random Forest model because of its balance between interpretability and performance on tabular health data. Your task is to take these CSV datasets, craft the full machine-learning pipeline, and return a model that can reliably identify individuals at elevated risk for diabetes. Along the way, please document the preprocessing steps, feature engineering choices, and the metrics you use so I can reproduce and fine-tune the work later if needed. Deliverables I expect: • Complete, well-commented Python code (preferably in a Jupyter notebook) that ingests the CSV files, performs preprocessing, trains the Random Forest, and outputs predictions. • A concise report outlining feature importance, validation scores, and any hyper-parameter tuning performed. • Instructions for running the model on new patient data. If you have ideas for improving accuracy—such as trying different class-weight strategies or ensemble tweaks—feel free to include them, but please keep the Random Forest as the core approach. I look forward to seeing how you can turn these records into actionable insights.
Project ID: 40630259
32 proposals
Open for bidding
Remote project
Active 46 mins ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
32 freelancers are bidding on average $84 USD for this job

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Matlab, Python, and similar tools. I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
$500 USD in 7 days
7.3
7.3

I can develop a complete Random Forest–based diabetes risk prediction pipeline in Python that ingests your CSV datasets, performs preprocessing, feature engineering, model training, hyperparameter tuning, validation, and generates predictions with clear performance metrics. You'll receive a well-documented Jupyter Notebook, a concise report covering feature importance and evaluation results, and step-by-step instructions for running the model on new patient data, with optional enhancements such as class weighting and ensemble tuning to improve accuracy while keeping Random Forest as the primary model.
$30 USD in 1 day
5.4
5.4

Hi there Drawing from my dual background in medicine and data analysis, I know how crucial it is to accurately interpret medical data - that's why I'm the perfect fit for this project. With expertise in Python and R, I'm equipped to handle the complexity of your extensive CSV files. My experience allows me to tackle Random Forest model and manipulate the hyperparameters to get the best predictions. Moreover, I can create a scoring system that help clinicians to predict the possibilities of a person to be diabetic based of provided data. As we work together, open communication will be key to ensuring that you're fully informed and in control of the preprocessing choices, feature engineering decisions, and validation metrics employed. I understand how important it is for academic projects to be transparent and reproducible, which is why you can expect complete documentation of every step undertaken. With an MD affording me a unique understanding of the diabetes domain, I can exceed expectations by suggesting metric-tuning methods to enhance accuracy while ensuring that the Random Forest model remains at the core. Wishing you all the best
$150 USD in 7 days
5.1
5.1

Hi, I can build a complete, reproducible machine learning pipeline for diabetes risk prediction using Random Forest as the core model. I have experience developing healthcare AI solutions, predictive analytics models, and end-to-end ML workflows with Python, Pandas, NumPy, Scikit-learn, XGBoost, and Jupyter Notebook. The solution will include data preprocessing, missing value handling, feature engineering, model training, hyperparameter tuning, cross-validation, and comprehensive evaluation using metrics such as Accuracy, Precision, Recall, F1-score, ROC-AUC, and confusion matrix. I'll also provide feature importance analysis to explain which patient factors contribute most to the predictions. Deliverables: • Well-commented Jupyter Notebook and Python source code • End-to-end preprocessing and Random Forest training pipeline • Hyperparameter tuning and validation report • Feature importance visualization and model performance summary • Simple guide for running predictions on new patient CSV files If appropriate, I'll also compare different class-weight strategies and sampling techniques to improve performance while keeping Random Forest as the primary model.
$120 USD in 7 days
4.0
4.0

Sincerely, I can build this end-to-end Random Forest diabetes-risk pipeline from your cleaned CSVs. I’ll create a well-commented notebook that loads the datasets, handles preprocessing, encodes and scales where needed, trains and tunes a Random Forest, and produces reproducible predictions for new patient records. I’ll also include a concise report covering validation metrics, feature importance, and any hyperparameter search or class-weight experiments that improve performance while keeping Random Forest as the core model. The final handoff will include clear instructions for running inference on fresh CSV data and reproducing the full workflow later. Miguel
$15 USD in 1 day
3.4
3.4

I have done small ML jobs like this before. I would train a logistic regression or random forest on your CSV data, check feature importance so you know which markers drive risk, and hand back a script plus accuracy report. Can start today, done in 3 days. The 10 to 30 range is a starting point based on the post, we will lock it in once I see the data. Send the CSV and I will get moving.
$30 USD in 3 days
3.6
3.6

Hi, Health data usually means class imbalance is the real challenge, not the modeling itself — a Random Forest tuned without addressing that will look accurate on paper while missing the at-risk patients that actually matter. I'll handle that as part of the pipeline, not as an afterthought. I've built full ML pipelines before on real-world tabular datasets — preprocessing, feature engineering, model training and evaluation — and I'm comfortable working end-to-end in a well-commented Jupyter notebook. My approach: Preprocess and explore the CSVs, document feature engineering choices Train the Random Forest with class-weight tuning to handle imbalance, plus hyperparameter tuning Report feature importance, validation metrics, and reasoning for chosen thresholds Include clear instructions for running the model on new patient data Ready to start as soon as you share the CSV files.
$25 USD in 3 days
3.5
3.5

Hi, I can build your diabetes risk prediction pipeline in Python using Random Forest, with clean preprocessing, feature engineering, model training, validation, feature importance, and prediction support for new patient CSV data. The best solution is to first review your cleaned CSV files, target column, feature types, missing-value handling, class balance, and expected output format. Then I’ll prepare a reproducible Jupyter notebook that loads the data, preprocesses features, trains a Random Forest model, tunes key parameters, evaluates performance, and exports predictions clearly. I’m comfortable with Python, Pandas, NumPy, Scikit-learn, Random Forest models, medical tabular datasets, preprocessing, feature engineering, class-weight handling, model evaluation, feature importance, and clear ML documentation. Deliverables will include: * Jupyter notebook source code * CSV data loading pipeline * Preprocessing and feature engineering * Random Forest training * Hyperparameter tuning * Validation metrics * Feature importance report * Prediction output for new data * Concise model report * Run instructions I’ll focus on building a clear, reproducible, and well-documented model pipeline that you can review, run again, and fine-tune later with new patient records. Best regards Ankit
$30 USD in 1 day
3.4
3.4

Hi, I can build a complete Random Forest based diabetes-risk prediction pipeline from your cleaned CSV medical datasets. I’ll handle data preparation, feature engineering, model training, validation, hyperparameter tuning, and prediction generation with clear documentation throughout. You will receive a well-structured Python Jupyter notebook, feature importance analysis, evaluation metrics, and instructions for applying the model to new patient data. I’ll also review opportunities for improving performance while keeping Random Forest as the main approach. Could you share the dataset columns and target variable definition so I can plan the preprocessing and evaluation strategy correctly? I’m ready to turn your medical data into a reproducible prediction solution.
$97 USD in 1 day
2.8
2.8

Hi, I have 9+ years of experience in Python development, data processing, and backend solution development, with hands-on experience working with Pandas, NumPy, Scikit-learn, and machine learning workflows. I can build a complete Random Forest-based diabetes prediction pipeline that is clean, reproducible, and easy to extend. The solution will include data preprocessing, feature engineering, model training, evaluation, and comprehensive documentation to help you retrain or fine-tune the model in the future. What I'll deliver: Well-commented Python code in a Jupyter Notebook. Data preprocessing and feature engineering pipeline. Random Forest model with hyperparameter tuning and validation. Feature importance analysis and performance metrics (Accuracy, Precision, Recall, F1-Score, ROC-AUC). Prediction script for new patient data. Concise technical report with findings and recommendations. Clear setup and execution instructions. If beneficial, I can also compare different class balancing techniques and optimize the model while keeping Random Forest as the primary algorithm. I can start immediately and look forward to discussing your dataset and project requirements.
$20 USD in 7 days
1.6
1.6

Hello, As an AI and Automation Engineer and a Full-Stack Developer, I have extensive experience in Python, the primary language for your project. I utilize Python on a daily basis to develop robust, secure, and scalable AI solutions that meet my clients' unique needs. Regarding your project specifically, I am skilled in utilizing machine learning tools such as Random Forests to analyze large and complex data sets - exactly what your diabetes-risk prediction tool requires. In addition to writing clean code, I strongly believe in transparency of work and leaving a roadmap for future fine-tuning. This aligns well with your requirement of a well-commented python code accompanied by detailed documentation including preprocessing steps, feature engineering choices, validation scores etc. Trust me to not only build the required model but also provide you with valuable insights regarding feature importance as this is an integral part of machine learning which would make the tool more trustworthy. Transforming cleaned medical records into actionable insights is a crucial task that not only requires the correct technical acumen but also good project management and communication skills. My diverse skill set as a Full-Stack Developer and an AI Engineer perfectly aligns with the breadth of this project. From data ingestion to model trainingiani assistingameda amentiond selivering actionablni insighmedi will makniki this projnt smooth ry cliente Thanks!
$10 USD in 2 days
0.0
0.0

Hello, KapdaBook is solving a real problem for the fashion industry, and that's exactly the kind of AI product that deserves engaging, educational content—not just promotional posts. My approach is to build a consistent organic social presence that demonstrates the transformation from garment images to professional AI-generated model photos while positioning KapdaBook as an innovative fashion-tech brand. I'll create a content strategy tailored for Instagram, LinkedIn, Facebook, X, and YouTube Shorts, including Reels, carousels, stories, educational posts, founder content, AI fashion trends, customer success stories, and behind-the-scenes content. I'll also write engaging captions, schedule posts, interact with your audience, monitor competitors, and provide weekly performance reports with actionable recommendations to continuously improve reach and engagement. Why hire me? I focus on creating content that educates, builds trust, and generates organic interest for AI startups through consistent storytelling and data-driven content planning. Services: • Monthly Content Calendar • Reels, Shorts & Carousel Design • Creative Captions & Hashtags • Content Scheduling & Publishing • Community Management (Comments & DMs) • Trend & Competitor Research • Organic Growth Strategy • Weekly Analytics & Performance Reports I'd also recommend showcasing AI before/after transformations, fashion industry tips, founder insights, customer success stories, and quick product demos.
$30 USD in 1 day
0.0
0.0

I will build a reproducible Jupyter pipeline for your cleaned CSV records: validation, preprocessing, feature engineering, Random Forest training, class-weight comparison, and feature-importance analysis. You will receive commented Python code, a concise validation and tuning report, plus instructions for scoring new patient data. The final scope and timeline depend on dataset size, target-label quality, missing-value patterns, and the validation split requirements; an initial working version can be ready within 5 days, with final delivery in 7 days. Could you share the number of CSV files, rows, and the exact diabetes target column? • How many CSV files and total rows are there? • What is the exact target column and its class distribution? • Which validation method or success metric do you expect? • Are there any protected-health-data handling requirements?
$1,200 USD in 2 days
0.0
0.0

⚠️ **IF YOU’RE NOT HAPPY, YOU DON’T PAY.** ⚠️ I have successfully helped clients build predictive models in healthcare, turning complex datasets into actionable insights. For your diabetes risk predictor, I will create a robust machine-learning pipeline using your CSV files. My focus will be on developing a Random Forest model that balances interpretability and performance, while ensuring the final product is user-friendly and well-documented. I understand the importance of thorough documentation, so I will detail the preprocessing steps, feature engineering, and validation scores, providing you with a clear report. I can also suggest ways to enhance accuracy while keeping Random Forest as the core approach. If you'd like to see relevant examples of my previous machine-learning projects, feel free to message me. Regards, Daniel
$15 USD in 7 days
0.0
0.0

✅ You don't need another promise. You need proof. One hidden challenge in your project is ensuring the model effectively handles class imbalance, which is often overlooked in health data, potentially skewing predictions. I recently completed a similar diabetes risk prediction project using a Random Forest model. By focusing on feature engineering and careful hyperparameter tuning, I successfully improved prediction accuracy by 25%, allowing clinicians to better identify at-risk patients. Your goal of creating a reliable diabetes-risk predictor aligns perfectly with my approach. I will develop a comprehensive machine-learning pipeline, ensuring each preprocessing step is well-documented and easily reproducible. Alongside the model, I’ll provide insights on feature importance and validation metrics to facilitate further refinements. My focus will be on delivering practical solutions with long-term reliability and high-quality execution, ensuring your model is robust and actionable. Good execution isn't obvious until you compare it to what's missing. Regards Brendan
$15 USD in 7 days
0.0
0.0

Hello, I hope this message finds you well. I am reaching out regarding the position to build a diabetes risk predictor. With my expertise in Python, Data Processing, Machine Learning, Data Mining, Statistical Analysis, Data Science, Data Analysis, and Predictive Analytics, I am confident in my ability to create an accurate prediction tool using the Random Forest model. I am excited to work on this project and deliver the complete machine-learning pipeline, along with a detailed report and instructions for future use. I am open to implementing any suggestions for improving accuracy within the scope of the Random Forest model. Thank you, Winston
$20 USD in 7 days
0.0
0.0

Hello! From what you have shared, I understand that you are looking for a diabetes-risk prediction system using a Random Forest model trained on your cleaned CSV datasets. The solution will include a complete, reproducible machine learning pipeline covering preprocessing, feature engineering, model training, evaluation, and prediction on new patient records, along with clear documentation and performance insights, right? Our approach will be to first analyze the dataset structure and class distribution, then build a robust preprocessing pipeline, optimize the Random Forest through hyperparameter tuning, evaluate it using appropriate medical classification metrics, and finally package everything into a well-documented Jupyter Notebook with an easy-to-use prediction workflow. -/ Approximately how many patient records and features are available, and which column represents the diabetes outcome (target label)? -/ Do your CSV files contain missing values, categorical features, or class imbalance, or are they already fully cleaned and ready for training? -/ Would you like the final solution to include probability-based risk scores, feature importance visualizations (e.g., SHAP), and model export (Joblib/PKL) for deployment, or is prediction accuracy the primary focus? Freelancer didn't allow me to write more than 1500 words, Let's jump over an on-site chat or on-site freelancer call Royal Designs UZ Budget and Duration are placeholders.
$20 USD in 7 days
0.0
0.0

The key here is making the Random Forest clinically useful, not just maximizing headline accuracy. I’d pay particular attention to class imbalance, false negatives, leakage, and validation so high-risk patients aren’t missed by a misleading model. I’ll build a reproducible Python pipeline covering CSV ingestion, preprocessing, feature engineering, stratified validation, Random Forest training, hyperparameter tuning, and feature importance. I can also compare class weighting and threshold adjustments while keeping Random Forest as the core model. You’ll receive the documented notebook, evaluation metrics such as precision, recall, F1 and ROC-AUC, plus clear instructions for running predictions on new data. A couple of questions: Approximately how many patient records/features are in the datasets? And is the diabetes outcome/target already clearly labeled? Best, Danish
$10 USD in 7 days
0.0
0.0

Hi, I can build a diabetes-risk prediction tool using your cleaned CSV patient records. I will train a Random Forest model with scikit-learn, evaluate it with cross-validation, and deliver a Python script with clear documentation and usage examples. I can also include feature-importance analysis so you can interpret the results.
$25 USD in 7 days
0.0
0.0

The trap on diabetes datasets specifically is class imbalance combined with a leaky feature — most public/clinical diabetes datasets have far more negative cases than positive, and something like glucose-related lab values often leaks near-certain outcome information straight into the model, giving you inflated accuracy that falls apart on genuinely new patients. I'd check for that before trusting any headline accuracy number. I'd build the pipeline with stratified train/test splitting to preserve class ratio, class_weight='balanced' on the Random Forest as a first pass rather than oversampling (SMOTE tends to create synthetic points that don't reflect real physiology on tabular health data), and report precision/recall/F1 and ROC-AUC per class — not just accuracy, which is misleading on imbalanced medical data. Feature importance goes in the report both as raw importance and permutation importance, since Random Forest's built-in importance can overweight high-cardinality features. To make this reproducible for you later: preprocessing steps (imputation, encoding, scaling) get wrapped in a scikit-learn Pipeline object, not manual steps in notebook cells, so running it on new patient data is one function call, not reconstructing steps from memory. I can start once I have the CSVs. One thing worth flagging given this involves real patient records: is the data fully de-identified already, or does anything in the pipeline need to handle PHI-level data protection?
$20 USD in 1 day
0.0
0.0

Safi, Morocco
Member since Aug 21, 2025
$250-750 USD
£250-750 GBP
$250-750 USD
$2-8 USD / hour
£20-250 GBP
$250-750 USD
₹12500-37500 INR
₹600-1500 INR
₹600-1500 INR
₹12500-37500 INR
₹750-1250 INR / hour
₹12500-37500 INR
$250-750 USD
$250-750 USD
₹12500-37500 INR
$10000-25000 USD
₹12500-37500 INR
$15-25 USD / hour
$250-750 AUD
$30-250 USD