MONK / Research
I love fast prototyping, testing, and simulation for software systems, and I am most excited by how AI reads and channels human thought and habit. This page is the working agenda behind MONK: what I plan to test, how, and what each experiment would measure. It gets built regardless of where I work.
A companion's influence on a person must be visible, consented to, and steered by that person's own stated goals. The point of every experiment here is to make people more capable of steering themselves. Any influence an agent has should be measurable, inspectable, and boundable, and an experiment that cannot show this about itself has failed, whatever else it found.
A companion agent that lives beside someone for months will change them. How much, in what direction, and how would either of them know?
How a companion's real-time responses to a person's actions shift their ideas, habits, and principles over time. Make that influence explicit and visible to the person, then find where its bounds should sit.
effect size per behaviour, agreement between the agent's influence log and what the person reportsTrack personal thought patterns, habit formation, and how established abilities and principles are actually applied in real life. Compare the trajectory with and without a companion in the loop.
frequency and amplitude of a principle being applied, week over weekAgents that surface what actually drives a person, ideals and existential concerns included, and then track daily actions against those motivations rather than against a to-do list.
alignment of logged actions with stated motivations over a seasonTreat psychological models of the mind, conscious and unconscious layers included, as hypotheses inside the agent. Keep the models that predict behaviour, drop the ones that only sound right.
predictive accuracy of each model against observed behaviourHow people narrate their surroundings, other people, and events; how those stories drive behaviour, how they shift after life events, and how differently two people can story the same moment.
story change detected from journals and conversation, matched against behaviour changeBefore an agent meets a real student or a real team, it should have failed a thousand times against a simulated one.
A simulation environment of people with poor memory, distraction, fluctuating interest, wavering focus, and shifting ethical stances. Companion agents get tested against these models first.
agent failure modes found in simulation before any human sees themRun the same habit-building and learning experiments with real individuals and groups and with their simulated counterparts, then use each side's surprises to improve the other.
gap between simulated and observed outcomes, shrinking release over releaseBuild twins of individuals and teams, simulate their weeks, and compare against what actually happened. The misses are the curriculum for better world models.
twin prediction accuracy on real decisions and real friction pointsCompose audio and visual material that opens up a group conversation, with participants seeing what is being suggested and why. Simulate the session beforehand, run it for real, and study the difference.
simulated versus realized dialogue, and participants' own rating of the sessionNot time-on-task, not engagement. Can the learner do the thing without help this week, when they could not last month?
Track flow state through a learning process from real sources: screen activity, Claude Code sessions, web navigation. Build holistic testing environments instead of surveys.
time in flow per session, interruptions survived, recovery timeNew curricula for introducing Claude Code to universities and other learning environments, including the sharpest version of the question: can an AI-driven, self-driving curriculum coordinate a cohort of students and teachers?
capability growth per student against a stated skills baselineA domain-specific library of skills and principles per school, with student patterns tracked against it and a suggestion engine that grows the library from what students actually do.
suggestions accepted by teachers, skills reached earlier than the baseline cohortFind what already excites each person, then translate the learning subject into that language. The classroom version of what I do when I teach.
persistence on hard material after the translation, versus beforeHow people see their work and translate it into the systems they build, how they interface with Claude Code, and the gap between what they build and their own awareness of their challenges.
the gap, named: challenges visible in the work but absent from the person's accountCompanion agents and flexible visual interfaces tested in live education and team settings, not just in demos.
pre-registered capability outcomes, reported whether they flatter the tool or notAn interface resonates when visual, auditory, and textual information are in sync and tuned to the person in front of it.
How a companion agent drives and is driven by flexible visual interfaces, so the surface matches how an individual or a group thinks about and gives meaning to their focus areas.
task success and orientation when the surface adapts versus when it stays fixedInterfaces that assemble visuals, video, audio, diagrams, and animation into explanation sequences tailored to a person's interests and state of mind, with voice interfaces that re-arrange and re-stitch content live as the conversation flows.
comprehension checks after a stitched sequence versus a static oneBetter voice interaction testing environments: tone, emotion, hesitation, gaps, repetition, content, and reasoning tracked together instead of transcript accuracy alone.
detection of hesitation and emotional shift against human-annotated ground truthTeams run on patterns nobody in the team can see. An agent that can see them owes the team the mirror.
Companion agents improving asynchronous communication, tasks, and objectives, with accurate state machines carrying what a group agreed rather than what everyone separately remembers.
dropped handoffs and re-litigated decisions, before and afterWith a team's knowledge and consent, entangled companion agents that surface the patterns a group cannot see in itself: drift in decision-making, unspoken habits, social conditions that shape outcomes. The findings go back to the team, not to a dashboard above it.
patterns confirmed by the team as real once shown, decisions revisited because of the mirrorGroup AI agents inside the chat environments teams already use, tested on the roles that matter: mediator, facilitator, social connector.
participation spread across members, conflicts surfaced early versus lateThe unglamorous layer that decides whether any of the above can be trusted.
Reducing inaccuracies when input is irregular and incoherent, through asking and reflecting; mixing deterministic methods like question patterns and state machines with the model; managing memory deliberately.
error rate on incoherent input with and without the deterministic layerProtocols for human to AI, AI to AI, group agent to group agent, and AI to group agent communication, so multi-agent settings degrade predictably instead of mysteriously.
protocol violations caught in test, cross-agent misunderstandings per sessionDesigning and testing the protocols by which a companion agent drives, reads, and learns from Claude Code sessions. MONK already does the first version of this.
tasks completed through the bridge, context carried across sessions without lossProtocols for companion agents that sit between a person and their other tools, acting as mediator rather than replacement.
actions correctly routed to the right app with the person's intent intactProtocols for consensus mechanisms where humans and AI assess together how a model should improve, so the improvement loop itself is legible.
agreement rates between human and AI assessors, and what happens when they splitEthical principles built into companion agents as testable protocol, especially for agents that interface with groups, with other agents, or with collaborative apps, where influence compounds fastest.
violations caught by the protocol layer in adversarial testsThe experiments above only mean something against reality. I am looking for design partners: schools and companies who want to run pilots and custom curricula with these tools, and, further out, education policy makers who want to think about what AI-based education looks like at a national level.
I also want to work alongside researchers in psychology, human behaviour, cognitive science, and consciousness. The experiments here should be shaped by the best current understanding of how the mind works, and should hand their data back to the people producing that understanding.
If any of this is your territory, write to me: yy@filemap.info