MONK / Research
These experiments come out of a position I set out first in the MONK manifesto. If you have not read that, it is the shorter way in.
I love fast prototyping, testing, and simulation for software systems, and I am most excited by how AI reads and channels human thought and habit. This page is the working agenda behind MONK: what I plan to test, how, and what each experiment would measure. It gets built regardless of where I work.
A companion's influence on a person must be visible, consented to, and steered by that person's own stated goals. The point of every experiment here is to make people more capable of steering themselves. Any influence an agent has should be measurable, inspectable, and boundable, and an experiment that cannot show this about itself has failed, whatever else it found.
A companion agent that works beside someone for months changes them. I want to know how much, in what direction, and whether both sides can see it.
Test how companion AI agents introduce ideas, habits, and principles in response to real-time human actions, and make that influence visible and measurable rather than silent.
effect size per behaviour, agreement between the agent's influence log and what the person reportsTest a personal AI agent's ability to track thought patterns over a long period, build and track habits, and track how established abilities and principles are applied to real life, their frequency and amplitude. Compare that against how interaction with a companion agent changes them.
frequency and amplitude of a principle being applied, week over weekBuild agents that find deep motivations, ideals and existential concerns included, and track daily actions against those motivations.
alignment of logged actions with stated motivations over a seasonExperiment with various psychological models and models of the human mind to test their accuracy, including the conscious and unconscious layers.
predictive accuracy of each model against observed behaviourTest how humans model their surroundings, other people, and events; how those stories affect behaviour; how the stories change; how life events influence them; and how different people and groups assign different stories to the same thing.
story change detected from journals and conversation, matched against behaviour changeRun a lot of interactions in simulation to improve the model, then test it in real life, then use what actually happened to improve the simulation as well as the model.
Build a simulation environment that models human behaviour and mental limitations: simulated humans with poor memory, distractions, fluctuating interests, wavering focus, and shifting ethical stances. Test companion AI agents against these human models.
agent failure modes found in simulation before any human sees themTest companion AI agents with individuals and groups for habit building and learning, compare the results against the simulated humans and groups, and improve both the simulations and the agents.
gap between simulated and observed outcomes, shrinking release over releaseBuild digital twins of both individuals and groups, simulate them, and test reality against them. Build more accurate simulations and world models from what actually happened.
twin prediction accuracy on real decisions and real friction pointsGather individuals and provide audio and visual material to open up the conversation, with participants seeing what is suggested and why. Simulate the social conditions and dialogue first, compare against the realized session, and improve the simulation.
simulated versus realized dialogue, and participants' own rating of the sessionCan I see the learning happening from real sources rather than asking about it afterwards, and can a curriculum drive itself?
Track flow state through a learning process using several data sources: screen tracking, Claude Code sessions, web navigation. Build holistic testing environments.
time in flow per session, interruptions survived, recovery timeExplore new curricula for introducing Claude Code to universities and other education environments, including automated curricula: can an AI-based self-driving curriculum effectively drive and coordinate a group of students and teachers?
capability growth per student against a stated skills baselineBuild and track a domain-specific skills and principles library for schools, track student patterns, principles, and methods against it, and expand the database into a suggestion engine driven by student actions and behaviour.
suggestions accepted by teachers, skills reached earlier than the baseline cohortFind the core excitements of individuals and translate learning subjects into subjects they already find exciting.
persistence on hard material after the translation, versus beforeTest how people see their work and focus, and how that translates into the systems they build. Track how they interface with Claude Code and how it resonates with their lives, and track the gap between what they build and their awareness of their own challenges.
the gap, named: challenges visible in the work but absent from the person's accountTest companion AI agents and visual user interfaces in real-life education and team settings.
pre-registered capability outcomes, reported whether they flatter the tool or notImproving the interface so it matches how individuals and groups think about and give meaning to their focus areas.
Test how companion AI agents sync with flexible visual user interfaces, and improve those interfaces to match how individuals and groups think about and give meaning to their focus areas.
task success and orientation when the surface adapts versus when it stays fixedBuild user interfaces that stitch visuals, videos, audio explanations, diagrams, and animations, and test how effectively they translate ideas. Test sequences tailored to individual interests and state of mind, and modify the voice interface to re-arrange and re-stitch content as the conversation flows.
comprehension checks after a stitched sequence versus a static oneBuild more accurate voice interaction testing environments: tone of voice, emotion tracking, hesitation, gaps, repetition, content, and reasoning.
detection of hesitation and emotional shift against human-annotated ground truthTracking the unconscious patterns a team runs on, how they drift, and how they affect decision-making, and giving that back to the team.
Test how companion AI agents can improve asynchronous human communication, tasks, and objectives with accurate state machines.
dropped handoffs and re-litigated decisions, before and afterWith the team's knowledge and consent, track unconscious patterns in groups and teams: how they change and drift, and how they affect decision-making. Use entangled companion agents to track social conditions and behaviour, and give the findings back to the team.
patterns confirmed by the team as real once shown, decisions revisited because of the mirrorTest group AI agents inside the group chat environments teams already use, on mediator, facilitator, and social connector abilities.
participation spread across members, conflicts surfaced early versus lateThe work underneath: irregular input, memory, deterministic methods, and the protocols between humans, agents, and group agents.
Reduce inaccuracies when irregular, incoherent input is generated, through asking and reflecting back. Mix in deterministic methods such as question patterns and state machines, and manage memory deliberately.
error rate on incoherent input with and without the deterministic layerBuild protocols for human to AI, AI to AI, group agent to group agent, and AI to group agent communication.
protocol violations caught in test, cross-agent misunderstandings per sessionDesign protocols and test how companion AI agents interface with Claude Code. MONK already does the first version of this.
tasks completed through the bridge, context carried across sessions without lossDesign protocols and test how companion AI agents interface with other apps as mediators.
actions correctly routed to the right app with the person's intent intactBuild protocols for consensus mechanisms on how to improve models through human and AI assessment processes.
agreement rates between human and AI assessors, and what happens when they splitBuild ethical principles and protocols into companion AI agents, especially agents that interface with groups, with group AI agents, or with other people's apps and collaborative tools, where influence compounds fastest.
violations caught by the protocol layer in adversarial testsThe experiments above only mean something against reality. I am looking for design partners: schools and companies who want to run pilots and custom curricula with these tools, and, further out, education policy makers who want to think about what AI-based education looks like at a national level.
I also want to work alongside researchers in psychology, human behaviour, cognitive science, and consciousness. The experiments here should be shaped by the best current understanding of how the mind works, and should hand their data back to the people producing that understanding.
If any of this is your territory, write to me: yy@filemap.info
A companion agent is only half of it. The other half is the visual tool a person is working inside, and most of what I am building sits in the join between the two. FileMap's blog is where I write about that side: what a spatial interface is for, and what changes when an agent can see what you are looking at.
The manifesto is the position. The work is what has been built so far.