Home/ Work/ Meta · Avatars Store
Meta · Avatars Store · Avatars 2.0

What Does Quality Even Mean Here?

Measuring what didn't have a measure — what "better" means when there's no benchmark, no precedent, and no agreement. Constructing, evaluating and benchmarking quality metrics for a billion digital identities.

My RoleLead researcher, Avatars Store team
MethodsCompetitive research, global benchmark surveys, severity interviews
OutcomeA shared definition of what to measure — informing quality initiatives and improving the Store over time through benchmarking
1B
Digital identities in scope
1,800
Users surveyed · 6 markets
5
EQR quality dimensions
20
Severity interviews
Act I The Meta Mockery

Meta Connect 2022. Meta Avatars became a global punchline.

Mark Zuckerberg posted a selfie from inside Horizon Worlds. The internet did the rest. The avatar was called "soulless," "dead-eyed," "basically a Wii Mii." Within days he was publicly promising a graphics upgrade. The headlines wrote themselves.

Press · The reception
PYMNTS

In Horizon Worlds, the Emperor Has No Legs

The avatars couldn't stand up — literally.

Forbes

Zuckerberg Promises Graphics Boost After Metaverse Mocking

A public promise to fix the way it looked.

Creative Bloq

Zuckerberg's Metaverse Avatars Now Look Slightly Less Ridiculous

The verdict online

"Soulless." "Dead-eyed." "Basically a Wii Mii."

The comparison nobody at Meta wanted.

The infamous "Zuckerberg avatar" meme crystallized a growing perception —the avatars were low-quality, uncanny, and out of touch.

Meta Avatars 1.0 — legless torsos floating in a stylized park, the version mocked globally.
Meta Avatars 1.0 · 2022Meta Avatars 1.0 — the version that was mocked globally. But "better graphics" is a dangerously vague north star when you're rebuilding digital identity from the ground up. And the only thing worse than launching a bad avatar was launching a "better" one that everyone hated.

When a billion people have collectively laughed at your product, "improvement" isn't just a design challenge — it's a credibility reconstruction.

Every design decision carried the weight of that first impression. Ship the wrong thing, and you don't just get bad reviews — you confirm the world's suspicion that you never understood the problem in the first place. The entire Avatars org faced three intersecting pressures.

Executive pressure
Leadership needed proof that quality was improving — not just anecdotes.
Global stakes
Perception varied dramatically by market — some markets had zero adoption.
No baseline
No shared definition of "quality" for the Store, let alone scores to track quarter-over-quarter.
Act II Constructing Measures
The mandate

My mandate was simple: make the Avatars Store better. Simple, except no one could define what "better" meant.

I was the lead researcher for the Avatars Store — where people created their avatars by choosing their face, hairstyle, clothing, and accessories. Yet the team had no benchmark for quality, no shared vocabulary, and no way to measure whether the experience was actually improving.

Rather than starting with new research, I first synthesized years of usability studies, past research, app reviews, and competitive products. The reviews were particularly revealing: users naturally compared Meta Avatars with other avatar systems, giving us a clear signal for what quality meant from their perspective.

Aligning to a shared language. Around the same time, the Avatars organization introduced an Experiential Quality Rubric (EQR) built around five dimensions — Performance, Usability, Trust, Craft, and Delight. I deliberately aligned my measurement framework with the EQR so every Avatar product could speak the same language internally and externally, while tailoring those broad dimensions into measurable signals for the unique context of the Avatars Store.

Strategic research questions

  1. How do we operationalize the EQR so that "quality" becomes measurable and actionable for the Avatars Store?
  2. What is the baseline quality across each EQR dimension globally, and where are the largest market-level gaps?
  3. How do platform-specific expectations shape users' perception of quality in the Avatars Store?
  4. Which dimensions have the greatest impact on users' perception of quality?
  5. Which issues most strongly drive low scores, and what is their impact on users' experience?

A three-phase sequential mixed-methods study

PhaseMethodPurposeN / Scope
Phase 1Analytics Audit + Heuristic EvaluationIdentify quality signals and operationalize them into survey dimensionsExpert review + behavioral data
Phase 2Global Quantitative SurveyBenchmark quality scores per dimension; test for significant differences across markets (ANOVA)n = 1,800 users across 6 countries
Phase 3Semi-Structured InterviewsExplain why scores were low; assess severity and emotional impactn = 20 participants

Operationalizing the EQR framework. The critical first step was translating the abstract EQR framework into Store-specific measures. Rather than relying on assumptions, I triangulated three sources of evidence:

  • Heuristic evaluation + analytics with Data Science to identify observable behavioral signals.
  • Comparative interviews where participants evaluated avatars they had created across Meta, Roblox, Zepeto, and other platforms. Comparing products revealed evaluation criteria far more naturally than asking users to rate Meta in isolation.
  • Cross-cultural research synthesis, particularly work from Korea and Japan, where avatar creation was already an established social behavior — preventing us from importing e-commerce assumptions into an identity product.

Operationalizing the EQR for the Store

Each abstract dimension became a Store-specific definition tied to an observable behavioral signal. Stakeholders trusted the scores because they mapped to things they could see in the data.

01
PerformanceSpeed, stability, and responsiveness of the Store

Signal: load latency correlated with session abandonment.

02
UsabilityEase of navigation, selection, and completion

Signal: high browse-to-save friction; time-to-completion outliers.

03
Trust & InclusivityAbility to create an avatar that reflects one's identity

Signal: demographic engagement gaps; support ticket themes.

04
DesignVisual fidelity, polish, and aesthetic appeal

Signal: low save rates on facial features; "uncanny" heuristic flags.

05
FunEnjoyment, delight, and self-expression pleasure

The assumption to challenge: treated as "not a must-have for a Store since it's a utility feature." I disagreed — and tested it.

Key strategic decision — was "Fun" a nice-to-have or a must-have?

The room's position

Many teams viewed Fun as optional because the product was a "Store" — a utility feature, secondary to the "serious" quality dimensions.

My position

Creating an avatar isn't a checkout flow — it's an identity experience. Prior research from APAC suggested delight and self-expression were core drivers of engagement, so I retained Fun as a measurable dimension and tested whether it predicted overall quality.

Operationalize to challenge assumptions. Building the survey from analytics and heuristics first meant every question was tied to an observable product behavior — even the ones designed to test what everyone took for granted.

Act III Quantitative · Global Benchmarking Survey

Designing a benchmark that could tell us how much, where, and whether we were improving

Once we finalized what to measure, I focused on designing a benchmark tool that could be used repeatedly to tell us how much, where, and whether we were improving. Why surveys, why three platforms (Facebook, Instagram, and Oculus), and why global:

N1,800 active Avatars users across 6 priority markets (US, UK, Brazil, Germany, India, South Korea)
InstrumentLikert scales for each EQR dimension, plus overall satisfaction and open-ended feedback
AnalysisDescriptive statistics per dimension; one-way ANOVA by country; post-hoc Tukey HSD for pairwise comparisons

The benchmark gave us our first global picture of quality — and one result immediately stood out. Fun showed one of the strongest relationships with overall quality, validating the earlier hypothesis that avatar creation behaves more like identity expression than a transactional store experience.

At the same time, Trust & Inclusivity remained among the lowest-rated dimensions in several markets, particularly on Facebook and Instagram. ANOVA by country turned diversity and inclusivity from a subjective conversation into a quantified quality deficit — which made it easier to prioritize against competing roadmap items. The survey showed where quality broke down, but not why these dimensions mattered — or what teams should fix first.

EQR benchmark · quality perception across markets and releases
EQR APAC NORAM LATAM EMEA ANOVA by country turned a subjective debate — into a quantified, prioritizable quality deficit.
Baseline
Re-measured
Illustrative — directional pattern, not exact values
Act IV Qualitative · Severity Interviews

The next question was why

I recruited participants from the lowest-scoring segments — not simply to validate the survey, but to understand the severity, emotional impact, and root causes behind the numbers. This mattered especially for dimensions like Trust & Inclusivity, where a one-point difference on a Likert scale could represent anything from a minor annoyance to feeling fundamentally misrepresented.

Three findings fundamentally changed how we interpreted the benchmark.

01

Performance wasn't just about speed — it was about expectations.

The survey suggested Facebook and Instagram users were less satisfied with performance than VR users, despite objectively slower load times in VR. The interviews explained the contradiction. VR users expected immersive experiences to take longer and were willing to wait. Mobile users expected instant feedback. The same delay felt acceptable in one context and frustrating in another — because the expectation was different.

02

Inclusivity was a content problem, not just a representation problem.

Low Trust & Inclusivity scores weren't driven solely by missing skin tones or body types. Across several markets, people struggled to find hairstyles, clothing, and accessories that reflected their local culture and identity. Users didn't simply want more options — they wanted options that felt relevant to where they lived.

03

Fun wasn't a bonus — it was the experience.

The strongest predictor from the survey came into focus during the interviews. Participants consistently described trying on outfits, experimenting with different looks, and seeing themselves come to life as the most enjoyable part of creating an avatar. Unlike traditional e-commerce, users weren't trying to complete a purchase as quickly as possible. They were exploring their identity. The try-on experience wasn't decoration — it was the product.

Act V · Immediate Roadmap Shifts

The research changed not just what we measured, but what we built.

01

Delight over raw performance

Rather than prioritizing VR performance improvements, teams focused on reducing perceived waiting on mobile while investing in moments of delight during avatar creation.

Reframed the roadmap
02

Try-on redesigned as the product

The try-on experience was rebuilt with larger previews, richer animations, and expressive poses that encouraged exploration instead of rushing users toward completion.

Exploration over checkout
03

Regionally relevant merchandising

Trust & Inclusivity findings sparked a merchandising initiative to expand regionally relevant clothing and accessories — helping users see themselves reflected in the Store rather than a one-size-fits-all catalog.

Content, not just representation
04

Discovery built for exploration

Instead of optimizing for traditional e-commerce findability, the roadmap prioritized personalized recommendations that helped users uncover styles they didn't know to search for.

Matched exploratory behavior

For the first time, teams weren't debating whether the Store was "good." They could point to a specific quality dimension, a specific market, and a specific user behavior — and decide exactly what to improve.

The same Meta avatar rendered in Sept 2021, Dec 2022, Apr 2023, and Sept 2024 — fidelity and expressiveness climbing each release.
Build the instrument to last · one self, four releasesDesigning the survey for re-fielding from day one meant that six months later, we could validate improvement with confidence. A one-off study tells you what to fix; a repeatable benchmark proves you fixed it. The benchmark became the organization's shared language for quality — adopted across Avatar teams and re-measured every six months.
Strategic Impact

The benchmark became the org's shared language for quality

As quality perceptions improved, so did engagement, adoption, and ultimately users' willingness to spend on digital self-expression. The framework turned "is it good?" into a specific, prioritizable conversation — one metric, one market, one behavior at a time.

01Performance
02Usability
03Trust & Design
04Fun & Delight
Press coverage of Avatars 2.0 — six polished, expressive Meta avatars with detailed clothing and faces, headlined 'Meta Avatars come to life.'
The reception · 2024From "the emperor has no legs" to "Meta Avatars come to life." The same press that ran the mockery covered the 2.0 release — expressive, legged, and detailed enough that people were dressing a self worth dressing.
The outcome that mattered

From "the emperor has no legs" to a product measured not by opinions but by a shared definition of quality — the biggest outcome wasn't just a better Store. It was giving the organization a repeatable way to build, evaluate, and improve digital identity at global scale.

Note: metrics described directionally; exact figures held per NDA.

Next Study

When nobody uses the button that does everything

Optym · RouteMAX →

Get in touch

Let's think something through together.