The project at a glance
ConsultBae supported a healthcare AI client in building a smart-toilet health monitoring ecosystem. The product used computer vision and sensor analytics to read stool and urine patterns through a camera and sensor setup installed inside a toilet commode, connected to a mobile application that turned those readings into everyday health insights.
Our role was the dataset. We built the annotated data that trained the models behind digestive health monitoring, hydration analysis, menstrual indicators, urinary irregularities, and preventive wellness alerts. The work combined controlled biological data collection, precise medical annotation, and a multi-stage quality process built for an unusually ambiguous visual environment.
The challenge
A smart toilet is a difficult place to teach a model to see. The subject matter varies enormously from person to person and day to day, the environment is wet and reflective, and the data is about as sensitive as data gets. Four challenges shaped the entire project.
The first was biological variability. Diet, hydration, medication, and menstrual phase all change how a sample looks, so a single fixed rule could never cover the range. We built adaptive annotation guidelines that told annotators how to label consistently across those very different biological responses.
The second was rare edge cases. Blood traces, severe constipation, heavy diarrhea, mucus-heavy samples, and menstrual overlap are exactly the situations the product most needs to get right, and exactly the ones that appear least often. We used special sourcing and dedicated review workflows so these cases were captured and validated rather than averaged away.
The third was medical sensitivity. This is intimate health data, so contributor consent, anonymization, restricted access, and secure handling were not add-ons. They were built into the collection process from the start.
The fourth was visual ambiguity. Boundaries between stool, urine, tissue, and the toilet bowl itself are rarely clean. Multi-pass QA was used to verify those boundaries and catch tissue overlap, shadows, and abnormal color patterns before anything reached the client.
Controlled-environment data collection
Visual detection alone would not have been enough, because two samples that look alike can mean very different things depending on what the body was doing. So we designed controlled biological response studies, logging participant diet, hydration, sleep, bowel-movement timing, and menstrual phase. This helped the client understand how different body types react under different conditions and pushed model performance well beyond surface-level image detection.
| Observation group | Purpose and expected learning |
|---|---|
| High-fiber diet | Analyze constipation improvement and stool consistency changes |
| Low hydration | Monitor urine concentration and dehydration indicators |
| High liquid intake | Track urine dilution, transparency, and frequency patterns |
| Spicy food intake | Observe digestive irregularity and diarrhea-related patterns |
| Dairy consumption | Track lactose sensitivity and digestion response |
| Iron-rich diet | Simulate dark-stool appearance and reduce false alerts |
| Fasting condition | Understand metabolic response and low-volume bowel patterns |
Annotation scope
Each sample was annotated across several dimensions, not just outlined. Stool was captured with tight polygon boundaries and tagged for medical attributes, while urine, menstrual indicators, and non-biological artifacts each carried their own label sets.
| Dataset area | Labels and examples |
|---|---|
| Stool | Tight polygon boundary, Bristol number, color, texture, density, mucus, blood traces |
| Urine | Color grade, transparency, cloudiness, foam, sediment, hydration indicators |
| Menstrual indicators | Menstrual blood tag, period-onset indicators, recurring pattern support |
| Objects and artifacts | Tissue paper, skin tags, foreign particles, reflection, shadow, occlusion, water distortion |
The medical backbone of the stool labeling was the Bristol Stool Scale, classifying each sample from Type 1 through Type 7. That single consistent scale let the model connect what it saw to recognized clinical patterns, from the hard, low-frequency stool of constipation to the watery, repeated episodes of diarrhea.
In a wet, reflective bowl, the hard part was not seeing the sample. It was teaching the model what was not one.
The reflection and liquid-distortion problem
The single most technically stubborn part of the project was the environment itself. Toilet water reflections, lighting variation, liquid distortion, bubbles, and shadows all create segmentation traps. Left unhandled, a model reads a patch of water shine or a shadow as a biological region and raises a health alert for something that was never there. Solving that took a dedicated set of labels aimed squarely at what the model should ignore.
Reflection masking separated water shine from real biological regions.
Water-region tagging helped the model understand submerged or partially visible stool.
Shadow and occlusion labels reduced false positives caused by the geometry of the bowl.
Lighting-inconsistency tags improved model robustness across different washroom conditions.
The QA workflow
Because a wrong label in medical data is far more costly than a wrong label elsewhere, every sample passed through a five-stage validation process rather than a single review.
| QA stage | Validation process |
|---|---|
| Stage 1 | Annotation completeness and label coverage check |
| Stage 2 | Polygon precision and boundary tightness review |
| Stage 3 | Bristol number and medical attribute verification |
| Stage 4 | Reflection, occlusion, tissue, blood, and menstrual indicator audit |
| Stage 5 | Final AI-ready dataset consistency validation |
Privacy and compliance
Sensitive biological data was handled under strict controls throughout. Contributor consent was managed before any sample was collected, identity information was separated from dataset metadata, and annotation access and storage were restricted and secured. Handling practices were aligned with HIPAA and GDPR expectations for data of this kind, so that quality never came at the cost of privacy.
Business impact
The finished, AI-ready dataset improved the client's sensor accuracy, cut false positives from reflections and lighting, and supported a connected mobile experience that delivers continuous stool and urine health insights.
| Impact area | Outcome |
|---|---|
| Model development | Accelerated training for stool, urine, and health-indicator detection |
| Sensor accuracy | Improved performance across reflections, water distortion, and lighting variation |
| User experience | Enabled daily wellness summaries, hydration reminders, and digestive health reports |
| Preventive health | Supported constipation, diarrhea, menstrual, hydration, and blood-trace alerts |
More than the outputs, the project showed what it takes to run a sensitive healthcare dataset end to end: biological image annotation, controlled data collection, medical attribute tagging, and multi-stage QA, held together in an environment designed to confuse a model at every turn.
Sensitive data. Ambiguous environments. Handled properly.
If your AI depends on data that is complex, sensitive, or hard to label consistently, tell us the outcome you need and we will build the dataset behind it.
Talk to our AI data team