HumanFB Fish Observation Experiment

Continuous individual tracking, controlled stimulation, emergent-theory generation, and autonomous hypothesis testing in aquarium communities

Project: HumanFB
Document date: August 16, 2026
Status: Experimental design and theory-development companion
Initial organism class: Small aquarium fish, with guppies as the leading first implementation

Executive summary

The fish observation experiment is a concrete non-human implementation of HumanFB. A camera system continuously identifies and tracks individual fish in an aquarium, initially recording at least position and body heading versus time. Synchronized environmental sensors and stimulus controllers record what was presented, what was actually delivered, the state of the aquarium before presentation, and the individual and collective motion that followed. The resulting corpus is a Personal—or, more generally, Organismal—Stimulus–Response Graph in which each fish has a longitudinal identity, behavioral history, social context, and evolving response profile.

Small aquarium fish are attractive first subjects because they inhabit a bounded, continuously observable environment; respond to controllable visual, acoustic, hydrodynamic, nutritive, spatial, and social changes; and produce large quantities of low-cost time-series data. Guppies are particularly promising because their natural color and shape variation can assist individual recognition. A controlled enrollment station can extend the landmark, shape, and color approach demonstrated by Zion and colleagues in 2008, while modern multi-camera tracking can preserve those identities after the fish return to a freely swimming community.

The initial scientific target is not direct recognition of emotion. Position data support direct measurements of locomotion, spatial preference, approach, withdrawal, startle, recovery, schooling, social proximity, and habituation. Constructs such as behavioral activation, attraction, avoidance, vigilance, exploration, or feeding motivation are model-based interpretations. Claims about fear, pleasure, anxiety, or other subjective states require converging evidence and remain versioned hypotheses rather than replacements for the observed trajectories.

The most ambitious extension is an AI theory engine. The system need not begin with a human-supplied behavioral hypothesis. It can first learn the ordinary dynamics of the aquarium, discover recurrent individual and group states, identify reliable temporal relationships and deviations, express those observations as quantified propositions, generate alternative explanations, and propose experiments that distinguish among them. When stimuli are recorded and randomized, the engine can test causal response hypotheses. When no deliberate stimulus is present, it can still develop coherent observational theories about behavioral phenotypes, social influence, rhythms, latent states, and hidden environmental causes—provided it distinguishes correlation, prediction, and intervention.

The governing scientific sequence is:

immutable observation
        ↓
reproducible quantitative pattern
        ↓
candidate interpretation or mechanism
        ↓
falsifiable prediction declared in advance
        ↓
controlled observation or intervention
        ↓
support, revision, or rejection

The project succeeds when the system predicts held-out individual and collective responses better than simple historical baselines, preserves uncertainty and provenance, and produces at least some autonomously generated hypotheses whose prospective predictions survive independent testing.

1. Experimental thesis

The aquarium is a continuously measurable dynamical environment containing several coupled organisms. Every fish has internal state that is only partially observable, a history of prior events, relationships with other fish, and a stream of sensory input. Movement is one outward projection of that system.

HumanFB therefore treats the experiment as more than object tracking. It asks:

Given the current behavioral state of each fish, the current configuration of the group, recent history, environmental context, and an optional measured stimulus, what individual and collective trajectory is likely to occur next?

This formulation supports three progressively stronger modes of inquiry:

  1. Description: Identify recurring motion, spatial, and social patterns.
  2. Prediction: Forecast future positions, behavioral modes, and group changes from earlier evidence.
  3. Intervention: Select or randomize stimuli that distinguish among competing explanations.

Prediction is the bridge between observation and theory. A fluent explanation generated after seeing an event is inexpensive. A useful theory must say what should happen in a later, held-out interval under declared conditions.

2. Why guppies are a strong first subject

Guppies combine small size, active movement, visible sexual dimorphism, variable coloration, social behavior, and ready availability. Adult males often carry distinctive color patches and fin shapes that may serve as natural biometric features. Females are typically less visually distinctive, creating a useful harder case for appearance-based identity.

The 2008 Zion study used side-view images of Red-Blond guppies, a narrow imaging aquarium, a dark background, controlled illumination, contour landmarks, body and tail measurements, normalized color regions, and a Bayesian classifier. It reported approximately 90% sex-classification accuracy from shape, approximately 96% from color, and up to 98.8% from combined features when calibrated within an image set. Performance declined across growers and imaging sets, exposing a domain-shift problem that modern calibration, augmentation, representation learning, and multi-site validation can address. The paper is available as the original Zion et al. specification of the experiment.

The fish experiment extends that work in four directions:

Zebrafish or medaka remain valuable alternate subjects because established tracking systems have demonstrated recognition of unmarked individuals in collectives. The current generation of idtracker.ai uses learned representations of individual image fragments and is directly relevant to maintaining identities through crossings.

3. Measurement architecture

3.1 Aquarium and imaging geometry

The initial observation aquarium should be optimized for measurement rather than decorative complexity. A top camera provides the primary horizontal trajectory. A synchronized side camera supplies vertical position, a second view through crossings, and additional appearance information. A later stereo or multi-view configuration can reconstruct three-dimensional coordinates after correcting for glass and water refraction.

The preferred primary tracking image is a dark fish silhouette against a spatially uniform diffuse background. A transparent tank bottom, diffuser, and constant tracking backlight produce stable boundaries. Near-infrared tracking illumination and an IR-sensitive monochrome camera can separate the measurement channel from visible-light stimuli, provided the chosen wavelength is empirically verified as behaviorally inert for the species and remains continuously active.

The rig should include:

3.2 Enrollment station

A narrow side-view station creates a high-quality identity and morphology record for each fish before it enters the community. The station captures both sides where possible and records:

The original Zion features remain useful as interpretable outputs and baselines. Learned image embeddings complement rather than erase those measurements.

3.3 Edge and storage architecture

Continuous raw video is processed locally for time-critical tracking. The edge system maintains a rolling raw buffer, emits trajectories and quality flags, and preserves full-resolution clips around stimuli, anomalies, identity uncertainty, and hypothesis-test windows. Cloud or server resources maintain durable objects, experiment metadata, annotations, training jobs, model versions, and hypothesis records.

cameras and environmental sensors
               ↓
clocked edge acquisition and rolling raw buffer
               ↓
detection → identity → pose → trajectory → quality flags
               ↓
local event journal and stimulus-response clips
               ↓
object storage + time-series store + response graph
               ↓
feature learning + prediction + theory generation
               ↓
versioned model or proposed experiment returned to edge

Network connectivity is not part of the control loop. A network outage must not stop acquisition or corrupt stimulus timing.

4. Evidence layers

The system preserves a strict hierarchy between what the camera measured and what a model inferred.

Layer Examples Status
Raw evidence Video frames, sensor samples, feeder confirmation, vibration trace Immutable observation while retained
Tracking observation Fish identity candidate, position, heading, body mask, confidence Derived measurement
Kinematics Speed, acceleration, turning rate, path tortuosity Deterministic or statistical derivation
Behavioral event Freezing, darting, approach, withdrawal, schooling transition Operational classification
Response construct Behavioral activation, attraction, avoidance, startle, recovery Model-based interpretation
Latent state Vigilance, exploration, feeding motivation, social engagement Theory-dependent estimate
Subjective interpretation Fear-like, pleasure-like, anxiety-like state High-level hypothesis requiring converging evidence

An example record might state:

Observation:
Fish G07 accelerated from 1.8 to 12.4 cm/s within 180 ms,
turned 96 degrees away from the measured vibration source,
and remained within 2 cm of the opposite wall for 41 seconds.

Operational classification:
Rapid withdrawal followed by boundary-associated immobility.

Interpretation:
High-confidence startle response.
Moderate-confidence threat-like response.
Affective valence not identified from this event alone.

This separation permits new models to reinterpret old behavior without modifying the trajectory itself.

5. Core individual measurements

The minimum timestamped track contains:

timestamp
fish_id or candidate identity set
species_id
x, y, and optional z position
body heading
body mask and visible fraction
position confidence
identity confidence
occlusion and reflection flags
camera and calibration identifiers

From these observations, compute:

Individual identity is probabilistic. After an unresolved crossing, the system may retain several candidate assignments and resolve them from later views. It must not generate a precise but fictional personal history by forcing a premature identity choice.

6. Collective and social measurements

Community behavior is not reducible to the sum of independent tracks. The system derives:

Collective zebrafish motion has been separated into relatively polarized schooling and weakly polarized shoaling states using alignment, speed, and spacing. That work provides an empirical anchor for state discovery in the proposed system. See From Schooling to Shoaling.

7. Controlled stimulus program

7.1 Light

Record actual intensity, spectrum, spatial distribution, onset shape, and duration. Conditions may include bright-to-dim, dim-to-bright, localized light, spatial gradients, and stable-light shams. Illumination level and background color are distinct stimuli and should not be conflated.

7.2 Food and predictive cues

Use verified food delivery, feeder-only shams, different delivery locations, and cues that sometimes predict later food. Time since prior feeding is a baseline variable. Measurements include approach latency, arrival order, group compression, displacement, persistence at the feeding site, and learned response to the predictive cue.

7.3 Vibration and acoustic delivery

Replace manual tank tapping with a calibrated actuator and measure the delivered vibration with an accelerometer or hydrophone. Include controller-only and actuator-motion shams. Repeated stimuli allow measurement of habituation, sensitization, spontaneous recovery, and group propagation. Precisely triggered acoustic-startle habituation has been quantified in zebrafish, providing a useful methodological precedent. See the zebrafish startle-habituation experiment.

7.4 Social and visual presentation

Initially use a neighboring transparent compartment, removable barrier, screen, or standardized robotic model rather than repeatedly changing the permanent population. Present conspecific, heterospecific, empty, static, and motion-matched controls. Zebrafish studies have quantified attraction to conspecific form and biological motion through spatial choice, illustrating how visible social stimuli can be isolated. See Perceptual mechanisms of social affiliation in zebrafish.

7.5 Environmental and object stimuli

Additional controlled variables include water flow, object appearance, hiding structures, temperature within a narrow study range, filter-state transitions, and structured changes in accessible space. Every intervention receives an intended protocol and a measured-delivery record.

8. Trial and baseline protocol

Continuous monitoring establishes ordinary rhythms before deliberate trials begin. A suggested initial program is:

  1. One week of passive baseline and identity validation.
  2. One week of randomized light and stable-light sham trials.
  3. One week of food, feeder-cue, and feeder-only trials.
  4. One week of controlled vibration and mechanical shams.
  5. One week of visual/social presentations and matched controls.
  6. One week of held-out predictions, repeat exposures, and model-selected experiments.

A provisional acute-trial envelope is:

10 minutes       pre-stimulus baseline
30–120 seconds   presentation or delivery
10–20 minutes    response and recovery observation

The trial begins only after a declared baseline-stability criterion is met. Trial order, presentation side, active/sham status, and stimulus intensity are randomized within time-of-day blocks. The prediction model is frozen before presentation. Derived interpretation is frozen before the trial is added to the training corpus.

9. Theory 1: persistent individual behavioral phenotypes

Proposition: Fish exhibit stable individual differences in locomotion, exploration, spatial preference, social proximity, and stimulus responsiveness that persist beyond momentary context.

The system would first cluster behavior without using identity labels. It would then ask whether episodes from the same fish are more similar than episodes from different fish after controlling for time of day, sex, size, social composition, and environmental state.

Testable predictions include:

Failure to outperform context-only and population baselines would weaken the theory of stable individual phenotype.

10. Theory 2: context-gated response profiles

Proposition: The response to a stimulus depends substantially on pre-stimulus state and group configuration, not merely on stimulus identity.

The same food cue may produce different motion when a fish is already near the feeder, recently fed, socially displaced, or in a low-activity phase. The same vibration may propagate differently through a cohesive school and a dispersed shoal.

The theory predicts that models containing baseline velocity, position, local density, nearest neighbors, recent stimuli, and time of day will outperform a stimulus-only model. A stronger test holds the delivered stimulus constant and prospectively predicts response variation from baseline state.

11. Theory 3: persistent social influence networks

Proposition: Certain individuals reliably initiate or amplify collective transitions, and those influence relationships are directional rather than reducible to proximity.

Candidate observations include one fish changing heading before group-centroid motion, repeated following relationships, and asymmetric displacement at shared resources. The theory engine can construct a time-lagged influence graph and compare it with graphs produced from time-shuffled trajectories.

Prospective tests include:

The word leader should be reserved for a model with prospective predictive value, not assigned from a visually compelling anecdote.

12. Theory 4: collective state transitions

Proposition: The aquarium alternates among a small number of recurrent collective regimes with different response properties.

Possible regimes include dispersed exploration, cohesive schooling, feeding aggregation, low-motion rest, boundary concentration, and post-startle freezing. An unsupervised state model can discover candidate regimes from group speed, polarization, dispersion, zone occupancy, and interaction-network structure.

The theory becomes testable when the model predicts:

13. Theory 5: habituation, sensitization, and recovery are individual and social

Proposition: Repetition changes response magnitude, but the learning trajectory varies among individuals and depends on group behavior.

For calibrated vibration, light transitions, or object appearances, estimate initial response, habituation rate, asymptote, spontaneous recovery, and context-specific renewal. The system can test whether a fish habituates because its own recent history changed, because neighbors stopped responding, or both.

Competing models include:

individual-only learning
group-mediated learning
shared fatigue or time effect
stimulus-device drift
identity or tracking artifact

Randomized inter-stimulus intervals, sham sequences, and tests with altered group composition distinguish these explanations.

14. Theory 6: anticipatory behavior and learned environmental clocks

Proposition: Fish learn predictive regularities in feeding, lighting, room activity, and equipment cycles and begin responding before the nominal event.

The theory engine searches for trajectory changes that consistently precede recorded events. It then distinguishes clock-based anticipation from immediate sensory cues by varying event time, presenting the cue without the outcome, and delivering the outcome without the usual cue.

Testable hypotheses include:

15. Theory 7: spatial preference is relational, not merely geometric

Proposition: A fish does not simply prefer a fixed coordinate; it prefers relationships such as distance from a neighbor, access to cover, visibility of a stimulus, water flow, or position within the group.

A heat map may suggest that a fish prefers one corner. Alternative explanations include feeder location, reflection, current, light gradient, dominance avoidance, or a recurrent neighbor. The AI should generate these alternatives and request discriminating changes: rotate the feeder location, reverse flow, exchange wall treatments, move the visual stimulus, or alter group composition.

The strongest theory is invariant under irrelevant coordinate changes. If a preference follows the feeder after the feeder moves, it is feeder-relative rather than tank-coordinate-relative.

16. Theory 8: deviations from personal baseline reveal hidden state changes

Proposition: An individual's departure from its own expected trajectory distribution is more informative than departure from a population average.

The system learns an expected range for location, speed, social contact, feeding response, and circadian activity for each fish. It then identifies sustained multivariate deviations and asks whether they precede visible morphological changes, missed feeding, altered social status, or environmental problems.

This is an anomaly theory, not an automatic diagnosis. A prospective evaluation freezes anomaly alerts and later compares them with independently assigned outcomes. The principal metrics are detection lead time and false-alert rate.

17. AI theory generation without a supplied hypothesis

17.1 What “hypothesis-free” means

The system is never literally assumption-free. Camera geometry, tracked variables, time resolution, preprocessing, model family, and criteria for an interesting pattern all impose inductive biases. The intended claim is narrower:

The operator need not specify the biological relationship to be tested before observation begins. The AI can search the recorded state space for reproducible structure and convert selected structures into explicit, falsifiable hypotheses.

This is automated exploratory science followed by confirmatory testing, not proof generated from pattern recognition alone.

17.2 Discovery pipeline

The AI scientist operates in stages.

Stage A — Learn the ordinary world

Build predictive models of individual position, velocity, group structure, environmental cycles, and observation quality. The first goal is accurate next-interval prediction, not natural-language explanation.

Stage B — Discover recurrent motifs and states

Identify repeated trajectory motifs, collective regimes, pair relationships, change points, periodicities, and unusual events. Require persistence across days or cohorts rather than selecting only the strongest pattern in one interval.

Stage C — Express observations as propositions

Examples include:

G03 changes heading 0.6–1.1 seconds before 68% of high-polarization group turns.

G07 occupies the north-east zone primarily when G02 is within one body length;
the zone preference disappears when G02 is elsewhere.

Group dispersion begins declining approximately four minutes before scheduled feeding,
even on days with no visible operator entry during that interval.

Each proposition cites the exact intervals, analysis version, uncertainty, effect size, and negative examples.

Stage D — Generate competing explanations

For the apparent leading behavior of G03, alternatives might be:

Generating alternatives is as important as generating the preferred theory.

Stage E — Propose discriminating tests

The experiment planner ranks interventions by expected information gain, cost, recovery time, and whether the actuator can deliver them precisely. It freezes a prediction from each candidate theory before the trial.

Stage F — Update theory status

The system records support, contradiction, boundary condition, or inconclusive result. It does not rewrite the original theory after observing the outcome. A revision receives a new version and an explanation of what changed.

17.3 Discovery with recorded stimuli

When stimulus events and measured delivery are available, the AI can search for:

Randomization and shams allow the engine to progress from prediction to causal claims about the intervention.

17.4 Discovery without recorded stimuli

Continuous passive observation can still produce valuable theories:

Without measured interventions, the system should phrase mechanisms cautiously. It may conclude that one motion pattern predicts another, but not that the first caused the second. Unexplained residuals can motivate new sensors. For example, repeated simultaneous turns near a particular time may prompt installation of a hydrophone, room-motion detector, pump-state logger, or higher-resolution light sensor.

17.5 Protection against self-deception

An autonomous theory engine can generate thousands of plausible stories. The protocol therefore requires:

Natural-language coherence is never an evaluation metric. Prospective predictive performance and successful discrimination among alternatives are.

18. Example autonomously generated theories

The following are examples of what the system might propose after passive and stimulated observation.

Discovered pattern Candidate theory Discriminating test
One fish precedes most cohesive group turns Persistent directional social influence Change its position relative to a localized cue and test whether group turning follows fish or cue
Group compresses before normal feeding time Learned temporal anticipation Shift feeding time and separately suppress the ordinary feeder cue
Startle response is weaker when the group is cohesive Social buffering or shared directional certainty Deliver matched vibration during prospectively classified cohesive and dispersed states
One corner is occupied only when a particular neighbor is present Social relationship rather than fixed-place preference Move visual/social access while keeping tank geometry constant
Activity changes before measured light transition Sensitivity to an unlogged preparatory cue Log relay sound, fixture leakage, room entry, and electromagnetic or vibration events
One fish's food approach slows over several days before visible change Personal-baseline anomaly predicts later condition Prospectively freeze alerts and score lead time against independent outcomes
Two species segregate only after feeding Resource-mediated interspecies spacing Randomize food location and quantity while holding light and time constant
Repeated vibration response recovers after group membership changes Habituation has a social component Compare retained individuals with newly composed groups under matched stimulus histories

Each theory has a defined failure condition. For example, the influence theory fails if G03 does not improve prediction on held-out group turns, or if the apparent effect vanishes when leading-edge position is controlled.

19. Prediction targets and baselines

The system must predict quantities that can be scored without subjective adjudication.

Individual targets

Collective targets

Required baselines

A sophisticated model is useful only if it beats these baselines on held-out future data.

20. Data and hypothesis records

The core stores include:

Store Contents
Raw object store Video, audio, vibration, environmental sensor streams
Identity registry Enrollment images, embeddings, morphology, uncertainty, lineage
Trajectory store Timestamped positions, headings, masks, and quality
Event journal Stimulus commands, measured delivery, equipment events, annotations
Feature store Kinematics, spatial metrics, group metrics, learned representations
Model registry Training corpus, code, parameters, validation, artifacts
Hypothesis registry Observation, alternatives, prediction, proposed test, status
Experiment registry Randomization, protocol, actuator limits, outcomes, deviations

A hypothesis record should contain:

hypothesis_id and version
discovery dataset and excluded confirmation dataset
quantified observation
candidate mechanism
alternative explanations
falsifiable predictions
proposed discriminating experiment
pre-trial model snapshot
result and uncertainty
status: proposed, testing, supported, bounded, contradicted, or retired

21. Validation and exit criteria

Tracking

Stimulus alignment

Predictive modeling

Theory engine

22. Phased implementation

Phase 0 — Optical and identity calibration

Build the enrollment station and tracking aquarium. Establish geometric, color, and timing calibration. Reproduce the Zion-style sex and morphometric baseline using fish-held-out and grower-held-out evaluation.

Phase 1 — Continuous individual trajectories

Track a small single-species cohort with no deliberate stimuli. Validate identity and discover ordinary spatial, locomotor, and group structure.

Phase 2 — Controlled positive and sham stimuli

Add verified food, light, and vibration events. Establish clear response measurements and causal trial infrastructure.

Phase 3 — Longitudinal personal response models

Predict each fish's held-out responses using identity, baseline state, social context, and prior exposures.

Phase 4 — Autonomous observational discovery

Permit the AI to propose quantified patterns and candidate theories from a discovery interval. Reserve later intervals for prospective scoring.

Phase 5 — Active experiment selection

Allow the AI to choose among preauthorized stimulus protocols according to expected ability to distinguish competing theories. The device layer remains constrained to validated intensities and schedules.

Phase 6 — Mixed-species and naturalistic extension

Add structural complexity, multiple species, and longer-duration social interventions after the measurement and theory lineage are reliable in the simpler tank.

23. Larger ends of the experiment

The same infrastructure supports several goals.

Basic behavioral science

It provides unusually dense individual and group histories for studying individuality, social influence, collective dynamics, learning, anticipation, and context-dependent behavior.

Automated experimental discovery

It creates a bounded environment in which an AI can move from passive observation to theory generation and then to controlled testing. The aquarium is complex enough to contain emergent behavior but instrumentable enough to preserve causal timing.

Precision aquaculture and monitoring

Individual identity and personal baselines can support sexing, grading, growth measurement, breeding records, feeding optimization, anomaly detection, and early recognition of changed behavior.

HumanFB architecture validation

The experiment tests whether HumanFB depends on human language. If the architecture can learn useful individual profiles and generate successful predictions from nonverbal trajectories, it demonstrates that voice memos are one evidence channel rather than the essence of the system.

Comparative response modeling

Once several species adapters exist, the platform can compare which response constructs transfer across organisms and which remain species-specific. Position, latency, habituation, spatial choice, and social proximity provide a shared measurement vocabulary without assuming shared subjective experience.

Conclusion

The fish observation experiment converts a continuously visible aquarium community into a rigorous stimulus–response laboratory. Its foundation is modest: preserve the identity and position of each fish through time, log environmental context, and synchronize deliberate stimuli with the resulting trajectories. From that base, the system can estimate individual behavioral phenotypes, collective states, context-gated responses, social influence, anticipation, habituation, and personal-baseline anomalies.

Its most consequential feature is the theory engine. An AI can begin without a human-supplied biological hypothesis, learn the aquarium's ordinary dynamics, discover reproducible patterns, and propose coherent explanations. But a discovered pattern becomes a scientific theory only when the system states alternatives, declares a falsifiable prediction before seeing confirmation data, and survives a controlled or prospectively observed test.

That discipline makes the aquarium more than a source of attractive trajectory plots. It becomes a compact environment for testing whether AI can perform a complete epistemic loop: observe, compress, explain, predict, intervene, revise, and preserve the evidence behind every conclusion.