I translate human values and intent into behavioral specifications engineers can implement and evaluations can grade — on shipping agentic products at consumer scale, across voice, text, and wearable surfaces where the right behavior differs by modality.
That isn't a copy task or a prompt tweak. It's a specification problem: define the boundary precisely enough that it can be trained toward, graded against, and enforced at runtime. It's the seam most teams leave unowned — where model behavior, evaluation, permissions, and audit all touch — and it's the one craft I've carried across three eras.
Behavior design gets harder as systems leave the screen. A chat interface can show you what it understood. A pair of glasses cannot. A robot in a hallway cannot. The correction surface disappears exactly when the stakes rise.
Meta's own research names this: Habitat 3.0 and Socially Intelligent Robots benchmarks agents on social navigation — a robot giving way to a person in a narrow corridor. That is not a perception problem. It is a behavioral contract, expressed in motion.
I work on the version of that problem you can hear.
MFA, Writing for Performance (CalArts). Years directing how meaning lands live — timing, intent, the gap between what's said and what's understood.
Directed dialogue, IVR, and NLU at regulated scale — intents, entities, and recovery for tens of millions of interactions where a wrong turn had real cost.
Behavioral boundaries for agentic products — the taxonomies, gold sets, grading guidance, and safeguards that make those boundaries trainable and enforceable.
Voice behavior for wearables with no screen to fall back on.
Designed the “Stats” voice shortcut on Oakley Meta — a single-word command, no wake phrase, that returns a live stats overview mid-effort.
Bare keywords are the hardest activation problem in eyes-free audio. One word, spoken during exertion, has to stay distinguishable from ordinary speech under motion, wind, and elevated breathing — and the wearer has no screen to correct a false fire. It ships off by default.
Audition committee — Meta Voice 2.0. Screening and callbacks for Aria and Helio, two of Meta AI's system voices.
Authored the Voice input modality guidelines in Meta's public Horizon design system — the reference developers build against when they add voice to a spatial app.
Covers transparent activation, privacy controls, data ownership, cognitive load, error handling, and multimodal interaction. Eighteen defined terms, so that teams shipping voice share a vocabulary before they share a spec.
Designed the rules governing when the Quest assistant signals with an earcon and when it speaks — the sound-versus-speech decision for moving through the system.
A governed multi-agent orchestration system for high-stakes contract analysis, where I treat separation of authority as safety architecture — each agent operates under a written behavioral contract and an explicit list of actions it will not take.
Designed how the system decides when to answer, ask, defer, or refuse — grounded in the known unreliability of self-reported model confidence. The specific mechanism is proprietary and withheld.
Authored the taxonomy the system uses to flag concealment and error — each entry pairing recognition criteria with the required response and, deliberately, what the system must not do, including dual-use and boundary cases.
An independent audit layer distinct from the components doing the work, permission controls that bound each agent, and citation verification kept separate from the generative path — integrity checked by something that didn't produce the output.
The system distinguishes a true capability failure from a question that belongs to a human professional — and routes accordingly instead of guessing.
Designed and tuned enterprise AI telephony handling 2M+ calls per week for prescription refills and Rx coverage — where precision, compliance, and trustworthy error handling weren't features, they were the requirement.
Defined the KPI framework for conversational success — instrumenting where each journey began and ended, and the criteria separating a completed task from a drop-out.
Ran the intent / entity taxonomy as a live loop: analyzed real calls, authored intents, re-sampled each cycle to close coverage gaps and correct mislabeled entities — improving precision and recall release over release.
Built the platform's first Spanish-language intent libraries, improving recognition and reducing fallback for bilingual callers.
Implemented error-recovery and override flows to keep continuity during misrecognition.
Chat-to-IVR expansion; preserved intent across modalities and reduced fallback through improved intent handling.
Enterprise scope from the start: interaction models and system behaviors for contact-center software serving 28,000 employees and 6,000 physicians; bilingual IVR for AutoZone, Virgin Money, Globe Life, and GAP.
Training data and structured input patterns to improve multilingual NLU coverage and inclusivity.
Bilingual ASR / IVR consultancy for U.S. / LATAM; Spanish-native voice assistants and cross-border CxD talent matching.
I came to AI through theater and music, and Meta's Creative Audio AI team hired me for that combination as much as for the conversation design. Directing a performance and specifying a behavior are the same act at different distances: deciding what something does when no one is telling it what to do.
For twenty years I've worked the same problem in different clothing: how meaning is made, and who controls it. On stage it was performance. In the enterprise it was millions of calls a week where a misheard word had a cost. Now it's models — where the behavior a system exhibits is the product, and the specification of that behavior is the leverage. The craft has been called conversation design, then AI interaction design, and now — because the model itself became the surface — model behavior design. The name keeps moving; the problem doesn't.
I don't think the performer and the systems designer are two people. The instinct is the same: know exactly what should be said, what should be withheld, and when the silence is the more honest choice. That's dramaturgy, and it's also model behavior. The best people at this are creative people — which is why the field pays for taste, not just rigor.
What I keep finding is a seam no one wants to own, and it's where the trust actually lives. So I write the specs — the taxonomies, the gold sets, the grading guidance, the contracts — that turn "the model should behave well" into something an engineer can build and an eval can grade. And I build my own governed systems to keep the craft honest.
If you're building agentic or multimodal systems and the question of when a model should help, ask, or refuse is still unowned — that's my seam.