MONK  /  Research

The experiments I want to run.

I love fast prototyping, testing, and simulation for software systems, and I am most excited by how AI reads and channels human thought and habit. This page is the working agenda behind MONK: what I plan to test, how, and what each experiment would measure. It gets built regardless of where I work.

Yildirim Yazganarikan (Leo) Founder, FileMap · co-creator of MONK updated August 2026
The rule ahead of all of it

A companion's influence on a person must be visible, consented to, and steered by that person's own stated goals. The point of every experiment here is to make people more capable of steering themselves. Any influence an agent has should be measurable, inspectable, and boundable, and an experiment that cannot show this about itself has failed, whatever else it found.

01 · Influence

What a companion actually changes

A companion agent that lives beside someone for months will change them. How much, in what direction, and how would either of them know?

Influence, measured and bounded

How a companion's real-time responses to a person's actions shift their ideas, habits, and principles over time. Make that influence explicit and visible to the person, then find where its bounds should sit.

effect size per behaviour, agreement between the agent's influence log and what the person reports
Thought patterns over long time

Track personal thought patterns, habit formation, and how established abilities and principles are actually applied in real life. Compare the trajectory with and without a companion in the loop.

frequency and amplitude of a principle being applied, week over week
Deep motivations

Agents that surface what actually drives a person, ideals and existential concerns included, and then track daily actions against those motivations rather than against a to-do list.

alignment of logged actions with stated motivations over a season
Models of mind, under test

Treat psychological models of the mind, conscious and unconscious layers included, as hypotheses inside the agent. Keep the models that predict behaviour, drop the ones that only sound right.

predictive accuracy of each model against observed behaviour
The stories people run on

How people narrate their surroundings, other people, and events; how those stories drive behaviour, how they shift after life events, and how differently two people can story the same moment.

story change detected from journals and conversation, matched against behaviour change
02 · Simulation

Break a simulation, not a classroom

Before an agent meets a real student or a real team, it should have failed a thousand times against a simulated one.

Humans, modelled honestly

A simulation environment of people with poor memory, distraction, fluctuating interest, wavering focus, and shifting ethical stances. Companion agents get tested against these models first.

agent failure modes found in simulation before any human sees them
The sim-to-real loop

Run the same habit-building and learning experiments with real individuals and groups and with their simulated counterparts, then use each side's surprises to improve the other.

gap between simulated and observed outcomes, shrinking release over release
Digital twins of people and groups

Build twins of individuals and teams, simulate their weeks, and compare against what actually happened. The misses are the curriculum for better world models.

twin prediction accuracy on real decisions and real friction points
Facilitated sessions, simulated first

Compose audio and visual material that opens up a group conversation, with participants seeing what is being suggested and why. Simulate the session beforehand, run it for real, and study the difference.

simulated versus realized dialogue, and participants' own rating of the session
03 · Learning

Learning, measured as capability

Not time-on-task, not engagement. Can the learner do the thing without help this week, when they could not last month?

Flow, instrumented

Track flow state through a learning process from real sources: screen activity, Claude Code sessions, web navigation. Build holistic testing environments instead of surveys.

time in flow per session, interruptions survived, recovery time
Claude Code curricula

New curricula for introducing Claude Code to universities and other learning environments, including the sharpest version of the question: can an AI-driven, self-driving curriculum coordinate a cohort of students and teachers?

capability growth per student against a stated skills baseline
A living skills and principles library

A domain-specific library of skills and principles per school, with student patterns tracked against it and a suggestion engine that grows the library from what students actually do.

suggestions accepted by teachers, skills reached earlier than the baseline cohort
Core excitements

Find what already excites each person, then translate the learning subject into that language. The classroom version of what I do when I teach.

persistence on hard material after the translation, versus before
What they build versus what they face

How people see their work and translate it into the systems they build, how they interface with Claude Code, and the gap between what they build and their own awareness of their challenges.

the gap, named: challenges visible in the work but absent from the person's account
Real classrooms, real teams

Companion agents and flexible visual interfaces tested in live education and team settings, not just in demos.

pre-registered capability outcomes, reported whether they flatter the tool or not
04 · Interface

Interfaces that think with you

An interface resonates when visual, auditory, and textual information are in sync and tuned to the person in front of it.

Companion and canvas in sync

How a companion agent drives and is driven by flexible visual interfaces, so the surface matches how an individual or a group thinks about and gives meaning to their focus areas.

task success and orientation when the surface adapts versus when it stays fixed
Stitched explanation

Interfaces that assemble visuals, video, audio, diagrams, and animation into explanation sequences tailored to a person's interests and state of mind, with voice interfaces that re-arrange and re-stitch content live as the conversation flows.

comprehension checks after a stitched sequence versus a static one
Voice, tested honestly

Better voice interaction testing environments: tone, emotion, hesitation, gaps, repetition, content, and reasoning tracked together instead of transcript accuracy alone.

detection of hesitation and emotional shift against human-annotated ground truth
05 · Groups

Groups, seen clearly

Teams run on patterns nobody in the team can see. An agent that can see them owes the team the mirror.

Asynchronous work that holds

Companion agents improving asynchronous communication, tasks, and objectives, with accurate state machines carrying what a group agreed rather than what everyone separately remembers.

dropped handoffs and re-litigated decisions, before and after
A mirror for teams

With a team's knowledge and consent, entangled companion agents that surface the patterns a group cannot see in itself: drift in decision-making, unspoken habits, social conditions that shape outcomes. The findings go back to the team, not to a dashboard above it.

patterns confirmed by the team as real once shown, decisions revisited because of the mirror
Group agents in group chats

Group AI agents inside the chat environments teams already use, tested on the roles that matter: mediator, facilitator, social connector.

participation spread across members, conflicts surfaced early versus late
06 · Protocols

Protocols and plumbing

The unglamorous layer that decides whether any of the above can be trusted.

Irregular input, handled

Reducing inaccuracies when input is irregular and incoherent, through asking and reflecting; mixing deterministic methods like question patterns and state machines with the model; managing memory deliberately.

error rate on incoherent input with and without the deterministic layer
Communication protocols

Protocols for human to AI, AI to AI, group agent to group agent, and AI to group agent communication, so multi-agent settings degrade predictably instead of mysteriously.

protocol violations caught in test, cross-agent misunderstandings per session
Companions that interface with Claude Code

Designing and testing the protocols by which a companion agent drives, reads, and learns from Claude Code sessions. MONK already does the first version of this.

tasks completed through the bridge, context carried across sessions without loss
Companions as mediators to other apps

Protocols for companion agents that sit between a person and their other tools, acting as mediator rather than replacement.

actions correctly routed to the right app with the person's intent intact
Consensus on model improvement

Protocols for consensus mechanisms where humans and AI assess together how a model should improve, so the improvement loop itself is legible.

agreement rates between human and AI assessors, and what happens when they split
Ethics as protocol, not policy

Ethical principles built into companion agents as testable protocol, especially for agents that interface with groups, with other agents, or with collaborative apps, where influence compounds fastest.

violations caught by the protocol layer in adversarial tests

In the field

The experiments above only mean something against reality. I am looking for design partners: schools and companies who want to run pilots and custom curricula with these tools, and, further out, education policy makers who want to think about what AI-based education looks like at a national level.

I also want to work alongside researchers in psychology, human behaviour, cognitive science, and consciousness. The experiments here should be shaped by the best current understanding of how the mind works, and should hand their data back to the people producing that understanding.

If any of this is your territory, write to me: yy@filemap.info