Measuring what didn't have a measure — what "better" means when there's no benchmark, no precedent, and no agreement. Constructing, evaluating and benchmarking quality metrics for a billion digital identities.
Mark Zuckerberg posted a selfie from inside Horizon Worlds. The internet did the rest. The avatar was called "soulless," "dead-eyed," "basically a Wii Mii." Within days he was publicly promising a graphics upgrade. The headlines wrote themselves.
In Horizon Worlds, the Emperor Has No Legs
The avatars couldn't stand up — literally.
Zuckerberg Promises Graphics Boost After Metaverse Mocking
A public promise to fix the way it looked.
Zuckerberg's Metaverse Avatars Now Look Slightly Less Ridiculous
"Soulless." "Dead-eyed." "Basically a Wii Mii."
The comparison nobody at Meta wanted.
The infamous "Zuckerberg avatar" meme crystallized a growing perception —the avatars were low-quality, uncanny, and out of touch.
When a billion people have collectively laughed at your product, "improvement" isn't just a design challenge — it's a credibility reconstruction.
Every design decision carried the weight of that first impression. Ship the wrong thing, and you don't just get bad reviews — you confirm the world's suspicion that you never understood the problem in the first place. The entire Avatars org faced three intersecting pressures.
My mandate was simple: make the Avatars Store better. Simple, except no one could define what "better" meant.
I was the lead researcher for the Avatars Store — where people created their avatars by choosing their face, hairstyle, clothing, and accessories. Yet the team had no benchmark for quality, no shared vocabulary, and no way to measure whether the experience was actually improving.
Rather than starting with new research, I first synthesized years of usability studies, past research, app reviews, and competitive products. The reviews were particularly revealing: users naturally compared Meta Avatars with other avatar systems, giving us a clear signal for what quality meant from their perspective.
Aligning to a shared language. Around the same time, the Avatars organization introduced an Experiential Quality Rubric (EQR) built around five dimensions — Performance, Usability, Trust, Craft, and Delight. I deliberately aligned my measurement framework with the EQR so every Avatar product could speak the same language internally and externally, while tailoring those broad dimensions into measurable signals for the unique context of the Avatars Store.
| Phase | Method | Purpose | N / Scope |
|---|---|---|---|
| Phase 1 | Analytics Audit + Heuristic Evaluation | Identify quality signals and operationalize them into survey dimensions | Expert review + behavioral data |
| Phase 2 | Global Quantitative Survey | Benchmark quality scores per dimension; test for significant differences across markets (ANOVA) | n = 1,800 users across 6 countries |
| Phase 3 | Semi-Structured Interviews | Explain why scores were low; assess severity and emotional impact | n = 20 participants |
Operationalizing the EQR framework. The critical first step was translating the abstract EQR framework into Store-specific measures. Rather than relying on assumptions, I triangulated three sources of evidence:
Each abstract dimension became a Store-specific definition tied to an observable behavioral signal. Stakeholders trusted the scores because they mapped to things they could see in the data.
Signal: load latency correlated with session abandonment.
Signal: high browse-to-save friction; time-to-completion outliers.
Signal: demographic engagement gaps; support ticket themes.
Signal: low save rates on facial features; "uncanny" heuristic flags.
The assumption to challenge: treated as "not a must-have for a Store since it's a utility feature." I disagreed — and tested it.
The room's position
Many teams viewed Fun as optional because the product was a "Store" — a utility feature, secondary to the "serious" quality dimensions.
My position
Creating an avatar isn't a checkout flow — it's an identity experience. Prior research from APAC suggested delight and self-expression were core drivers of engagement, so I retained Fun as a measurable dimension and tested whether it predicted overall quality.
Operationalize to challenge assumptions. Building the survey from analytics and heuristics first meant every question was tied to an observable product behavior — even the ones designed to test what everyone took for granted.
Once we finalized what to measure, I focused on designing a benchmark tool that could be used repeatedly to tell us how much, where, and whether we were improving. Why surveys, why three platforms (Facebook, Instagram, and Oculus), and why global:
| N | 1,800 active Avatars users across 6 priority markets (US, UK, Brazil, Germany, India, South Korea) |
| Instrument | Likert scales for each EQR dimension, plus overall satisfaction and open-ended feedback |
| Analysis | Descriptive statistics per dimension; one-way ANOVA by country; post-hoc Tukey HSD for pairwise comparisons |
The benchmark gave us our first global picture of quality — and one result immediately stood out. Fun showed one of the strongest relationships with overall quality, validating the earlier hypothesis that avatar creation behaves more like identity expression than a transactional store experience.
At the same time, Trust & Inclusivity remained among the lowest-rated dimensions in several markets, particularly on Facebook and Instagram. ANOVA by country turned diversity and inclusivity from a subjective conversation into a quantified quality deficit — which made it easier to prioritize against competing roadmap items. The survey showed where quality broke down, but not why these dimensions mattered — or what teams should fix first.
I recruited participants from the lowest-scoring segments — not simply to validate the survey, but to understand the severity, emotional impact, and root causes behind the numbers. This mattered especially for dimensions like Trust & Inclusivity, where a one-point difference on a Likert scale could represent anything from a minor annoyance to feeling fundamentally misrepresented.
Three findings fundamentally changed how we interpreted the benchmark.
The survey suggested Facebook and Instagram users were less satisfied with performance than VR users, despite objectively slower load times in VR. The interviews explained the contradiction. VR users expected immersive experiences to take longer and were willing to wait. Mobile users expected instant feedback. The same delay felt acceptable in one context and frustrating in another — because the expectation was different.
Low Trust & Inclusivity scores weren't driven solely by missing skin tones or body types. Across several markets, people struggled to find hairstyles, clothing, and accessories that reflected their local culture and identity. Users didn't simply want more options — they wanted options that felt relevant to where they lived.
The strongest predictor from the survey came into focus during the interviews. Participants consistently described trying on outfits, experimenting with different looks, and seeing themselves come to life as the most enjoyable part of creating an avatar. Unlike traditional e-commerce, users weren't trying to complete a purchase as quickly as possible. They were exploring their identity. The try-on experience wasn't decoration — it was the product.
The research changed not just what we measured, but what we built.
Delight over raw performance
Rather than prioritizing VR performance improvements, teams focused on reducing perceived waiting on mobile while investing in moments of delight during avatar creation.
Reframed the roadmapTry-on redesigned as the product
The try-on experience was rebuilt with larger previews, richer animations, and expressive poses that encouraged exploration instead of rushing users toward completion.
Exploration over checkoutRegionally relevant merchandising
Trust & Inclusivity findings sparked a merchandising initiative to expand regionally relevant clothing and accessories — helping users see themselves reflected in the Store rather than a one-size-fits-all catalog.
Content, not just representationDiscovery built for exploration
Instead of optimizing for traditional e-commerce findability, the roadmap prioritized personalized recommendations that helped users uncover styles they didn't know to search for.
Matched exploratory behaviorFor the first time, teams weren't debating whether the Store was "good." They could point to a specific quality dimension, a specific market, and a specific user behavior — and decide exactly what to improve.
As quality perceptions improved, so did engagement, adoption, and ultimately users' willingness to spend on digital self-expression. The framework turned "is it good?" into a specific, prioritizable conversation — one metric, one market, one behavior at a time.
From "the emperor has no legs" to a product measured not by opinions but by a shared definition of quality — the biggest outcome wasn't just a better Store. It was giving the organization a repeatable way to build, evaluate, and improve digital identity at global scale.
Note: metrics described directionally; exact figures held per NDA.
Get in touch