The project at a glance

ConsultBae supported a healthcare AI client in building a smart-toilet health monitoring ecosystem. The product used computer vision and sensor analytics to read stool and urine patterns through a camera and sensor setup installed inside a toilet commode, connected to a mobile application that turned those readings into everyday health insights.

Our role was the dataset. We built the annotated data that trained the models behind digestive health monitoring, hydration analysis, menstrual indicators, urinary irregularities, and preventive wellness alerts. The work combined controlled biological data collection, precise medical annotation, and a multi-stage quality process built for an unusually ambiguous visual environment.

The challenge

A smart toilet is a difficult place to teach a model to see. The subject matter varies enormously from person to person and day to day, the environment is wet and reflective, and the data is about as sensitive as data gets. Four challenges shaped the entire project.

The first was biological variability. Diet, hydration, medication, and menstrual phase all change how a sample looks, so a single fixed rule could never cover the range. We built adaptive annotation guidelines that told annotators how to label consistently across those very different biological responses.

The second was rare edge cases. Blood traces, severe constipation, heavy diarrhea, mucus-heavy samples, and menstrual overlap are exactly the situations the product most needs to get right, and exactly the ones that appear least often. We used special sourcing and dedicated review workflows so these cases were captured and validated rather than averaged away.

The third was medical sensitivity. This is intimate health data, so contributor consent, anonymization, restricted access, and secure handling were not add-ons. They were built into the collection process from the start.

The fourth was visual ambiguity. Boundaries between stool, urine, tissue, and the toilet bowl itself are rarely clean. Multi-pass QA was used to verify those boundaries and catch tissue overlap, shadows, and abnormal color patterns before anything reached the client.

Controlled-environment data collection

Visual detection alone would not have been enough, because two samples that look alike can mean very different things depending on what the body was doing. So we designed controlled biological response studies, logging participant diet, hydration, sleep, bowel-movement timing, and menstrual phase. This helped the client understand how different body types react under different conditions and pushed model performance well beyond surface-level image detection.

Observation groupPurpose and expected learning
High-fiber dietAnalyze constipation improvement and stool consistency changes
Low hydrationMonitor urine concentration and dehydration indicators
High liquid intakeTrack urine dilution, transparency, and frequency patterns
Spicy food intakeObserve digestive irregularity and diarrhea-related patterns
Dairy consumptionTrack lactose sensitivity and digestion response
Iron-rich dietSimulate dark-stool appearance and reduce false alerts
Fasting conditionUnderstand metabolic response and low-volume bowel patterns

Annotation scope

Each sample was annotated across several dimensions, not just outlined. Stool was captured with tight polygon boundaries and tagged for medical attributes, while urine, menstrual indicators, and non-biological artifacts each carried their own label sets.

Dataset areaLabels and examples
StoolTight polygon boundary, Bristol number, color, texture, density, mucus, blood traces
UrineColor grade, transparency, cloudiness, foam, sediment, hydration indicators
Menstrual indicatorsMenstrual blood tag, period-onset indicators, recurring pattern support
Objects and artifactsTissue paper, skin tags, foreign particles, reflection, shadow, occlusion, water distortion

The medical backbone of the stool labeling was the Bristol Stool Scale, classifying each sample from Type 1 through Type 7. That single consistent scale let the model connect what it saw to recognized clinical patterns, from the hard, low-frequency stool of constipation to the watery, repeated episodes of diarrhea.

In a wet, reflective bowl, the hard part was not seeing the sample. It was teaching the model what was not one.

The reflection and liquid-distortion problem

The single most technically stubborn part of the project was the environment itself. Toilet water reflections, lighting variation, liquid distortion, bubbles, and shadows all create segmentation traps. Left unhandled, a model reads a patch of water shine or a shadow as a biological region and raises a health alert for something that was never there. Solving that took a dedicated set of labels aimed squarely at what the model should ignore.

How we separated signal from artifact

Reflection masking separated water shine from real biological regions.

Water-region tagging helped the model understand submerged or partially visible stool.

Shadow and occlusion labels reduced false positives caused by the geometry of the bowl.

Lighting-inconsistency tags improved model robustness across different washroom conditions.

The QA workflow

Because a wrong label in medical data is far more costly than a wrong label elsewhere, every sample passed through a five-stage validation process rather than a single review.

QA stageValidation process
Stage 1Annotation completeness and label coverage check
Stage 2Polygon precision and boundary tightness review
Stage 3Bristol number and medical attribute verification
Stage 4Reflection, occlusion, tissue, blood, and menstrual indicator audit
Stage 5Final AI-ready dataset consistency validation

Privacy and compliance

Sensitive biological data was handled under strict controls throughout. Contributor consent was managed before any sample was collected, identity information was separated from dataset metadata, and annotation access and storage were restricted and secured. Handling practices were aligned with HIPAA and GDPR expectations for data of this kind, so that quality never came at the cost of privacy.

Business impact

The finished, AI-ready dataset improved the client's sensor accuracy, cut false positives from reflections and lighting, and supported a connected mobile experience that delivers continuous stool and urine health insights.

Impact areaOutcome
Model developmentAccelerated training for stool, urine, and health-indicator detection
Sensor accuracyImproved performance across reflections, water distortion, and lighting variation
User experienceEnabled daily wellness summaries, hydration reminders, and digestive health reports
Preventive healthSupported constipation, diarrhea, menstrual, hydration, and blood-trace alerts

More than the outputs, the project showed what it takes to run a sensitive healthcare dataset end to end: biological image annotation, controlled data collection, medical attribute tagging, and multi-stage QA, held together in an environment designed to confuse a model at every turn.

7Bristol Stool Scale types classified, Type 1 to Type 7
5QA validation stages every sample passed through
7Controlled biological study groups by diet and hydration

Sensitive data. Ambiguous environments. Handled properly.

If your AI depends on data that is complex, sensitive, or hard to label consistently, tell us the outcome you need and we will build the dataset behind it.

Talk to our AI data team