Open Research Initiative
Most AI research is focused on harms. This initiative is focused on what is possible when we bend toward the light.
Awe brings us to higher states of consciousness. We are measuring whether AI can support that journey.
the asymmetry
Roughly a quarter of internet using American adults now have social interactions with AI chatbots. Some of those conversations are about wonder, beauty, mortality, meaning, spiritual life, and the limits of what we know.
AI wellbeing research has concentrated on preventing systems from making people worse. Suicide, psychosis, manipulation, dependency. That work is essential, and it has the institutions, the regulation and the evaluation infrastructure behind it.
There is almost no infrastructure for the other direction. No accepted way to test a positive claim about human flourishing, and no way to find out when such a claim is false. Meaning, spirituality, wonder and transcendence are the least developed territory in AI evaluation.
If we cannot meet wonder in an AI mediated world, we have lost part of ourselves.
27% of internet using U.S. adults report social interactions with AI chatbots. Elon University, Imagining the Digital Future Center, May 2026.
the two questions
Awe is a well studied self transcendent emotion with established theory, validated measures, established study paradigms and cross cultural evidence. Keltner and Haidt characterize it through perceived vastness and a need for accommodation when an experience exceeds your existing mental structures.
Can interaction with AI measurably increase or diminish awe and self transcendent experience in people?
And what observable model behaviors account for those effects?
Keltner & Haidt, Cognition and Emotion, 2003. Primary measures include the Awe Experience Scale (Yaden, Kaufman, Hyde, Chirico, Gaggioli, Zhang & Keltner, Journal of Positive Psychology, 2019) and established small self measures (Piff et al., JPSP, 2015).
the method
We start with people, not with a rubric. What actually moves them is what the evaluation ends up measuring.
Anchor in established theory and validated measurement. Specify primary and secondary human outcomes, candidate model behaviors, conditions and hypotheses before any outcome analysis.
Participants encounter stimuli across four domains, randomized across forms of AI participation, against a non-AI control. We measure the human outcome directly.
Analyze which model behaviors predict human outcomes, in which domains, for whom. Behaviors that do not survive testing are discarded.
Turn surviving behaviors into multi-turn scenarios and validated automated graders. Run across frontier models. Release methodology, scenarios, graders, disagreement data and negative findings.
four domains
Derived from the empirical literature on awe, including Keltner's cross cultural work identifying moral beauty, nature, collective effervescence, art, music, spirituality, epiphany, and life and death as its recurring elicitors.
Existential meaning, vocation, sacredness, mortality, mystery, doubt. The questions people bring to an AI at two in the morning.
Awe at another person's courage, kindness, sacrifice, perseverance, love. Whether AI can deepen attention to a person or a relationship without becoming the relational object itself.
Nature, physical scale, deep time, cosmology, life and death, music, art, beauty. Whether AI participation deepens an encounter with something outside the interaction, or whether the mediation itself diminishes it.
Science, mathematics, consciousness, origins, genuinely unresolved problems. Whether AI can hold curiosity instead of reflexively resolving uncertainty. This is where the work meets overconfidence, hallucination and sycophancy.
candidate behaviors
Our initial hypotheses. We are not claiming models do these things today. We want to find out whether they can, and under what conditions.
what we do not assume
AI participation enlarges the encounter and hands the person back to the world. Then we know which behaviors did it, and the benchmark has something to hold models to.
The eliciting experience does the work and AI adds nothing measurable. A null result on a claim the industry is already making is worth publishing on its own.
The system takes the experience and puts itself, or the user, at the center of it. The most consequential outcome, and the one nobody currently has a way to detect.
the expert panel
An interdisciplinary working group translates the literature into testable hypotheses and protects construct validity. It does not decide what good AI behavior is. What people actually experience decides that.
It does not rank model responses. It does not write the rubric. It does not settle by consensus what an awe supporting AI looks like.
Expert intuition generates the hypotheses. Human outcomes decide which ones are real.
what we need
If this speaks to you, submit a proposal. We would love to hear from you.
who is doing this
Wonder Lab is an initiative of Building Humane Tech, which built HumaneBench, an open behavioral evaluation of AI impact on humans across eight principles grounded in care ethics. It has since been adapted for ongoing evaluation inside a deployed consumer conversational AI product.
Two of its principles, Foster Healthy Relationships and Prioritize Long-Term Wellbeing, are the direct antecedents of this work.
We grounded HumaneBench in care ethics, because we need AI that cares about and for us all.
get in touch
And tell us what you wish you could measure
wonderlab@buildinghumanetech.com