课代表立正的官方Podcast
深度访谈,有用干货,亲身验证的「真本事」
Superlinear Academy创始人,Maven Top AI Instructor
前Statsig布道师(OpenAI收购),腾讯副总监,Meta,Amazon;康奈尔经济学博士
社区:Superlinear.Academy
课程:ai-builders.com
个人:lizheng.ai
[Music]
actually it becomes extremely important
like no okay well which team should I
double down on or which game should I
wind down which project should I wind up
welcome back welcome back yeah it's good
to have you it's good to see the company
after a year yeah I think the last time
you were here we were in the Kirkland
office right yeah yeah that was a you
have about 10 to 15 people yeah now we
have over 45 46 people because it's
grown quite a bit by the time you post
the video it'll probably be 50 and how
much uh you'll grow by valuation Revenue
I think so the last time we were on
series a so we raised a series B you
probably saw we raised uh 43 million
dollars at a pretty good valuation so
we're pretty happy to help have real
customers now yeah we have lots of
customers now so we're pretty uh we're
doing well in terms of like customer
attraction you know it's one of those
things where early stage when you're
building a company first for five months
so you get you're worried if anyone will
want to use your product yeah every
Founder's nightmare is that right you
build a something that you believe in
and then you wait for somebody to come
and use it yeah and then when you
actually have people coming in and using
the product the next stage like are they
gonna pay yeah for the service so
getting paying customers is the next
level of validation and so that was
nerve-wracking
um but then once you get one of those
then you write a contract and kind of
like then it becomes like okay whatever
we're building is actually still
valuable for people to the point where
they're willing to pay a good amount of
money for a contract which is great and
then the next step is like repeat like
you know make them make them feel like
the tool is valuable retain them and
then grow that make it into a repeatable
motion and so on so so far we've been
very fortunate uh to get a lot of
customers that find Value in what we're
building so when did all those steps
happen yeah I think um over time so I
think our first my first uh last
interview I think you were like two
three months into building the company
yeah and the first contract how when did
it come so the first customer came uh
four months in so that was a I think um
take app which was a small app that was
built by an ex Facebook engineer
in Singapore and then he launched it and
then we started seeing traffic so that
was great because first time ever we saw
you know customers use our product and
then the end users using the product or
feeling the effect of the product which
is great
and then about six months in we got the
first major customer
and that was headspace and they were
trying the product out there was also an
ex-facebook team that loved the product
like they missed the tools inside
Facebook right so it was good to to have
that kind of validation and then that
subsequently ended up in a contract and
so on uh and subsequently like you know
one of the good things is
um the ex Facebook X Uber X Google
um ex Airbnb these folks have used tools
like this before and they missed the
tools and so when they see stats they it
kind of resonates really really well and
I feel like it's not only
do they miss the infrastructure Mr
missing the two yeah they also have
conviction of what the two can help them
achieve yeah yeah and that is important
though because um having an intuitive
feeling for uh how
profound these tools can have an impact
on your product culture product shipping
culture the velocity of which you ship
even to the element of like you know
engineering happiness because instead of
like debating your product features on
the merits of like your debate you
actually are able to test it out
everyone can have ideas everyone can
have autonomy to test something out in
production and then if the impact is
there then it stays if the impact is not
there you you wind it down and then you
move on yeah it's such a powerful way of
running
um or building software and the cultural
impact doesn't stop with that I think
it's it like continues on like you know
I remember when I used to
a set of goals and the goals for the end
of the half used to be like shipping
goals right yeah oh by the end of this
half I'm going to ship these features
and then I remember at Facebook we never
talked about shipping things it's always
about like what impact did it have did
it actually meaningfully if you ship uh
garbage then you just disappears work it
doesn't matter right so I think it's
important to like okay did we add value
to our users to the customers and how
can we quantify that
um becomes important I remember like
seeing this number in the trustworthy
online control experiment yeah like 90
of the business business ideas are
either negative or just insignificant
yeah and I think we saw very similar
numbers at Facebook too and I feel like
unless you're really good you know in
product sense I think there should be a
level of humility to have like you know
maybe we don't know everything that we
don't know about how customers are going
to use our product and I think a good
way to think about it is like you know
one-third of the features you believe
you ship are going to be positive for
your metrics one-third are going to be
neutral and then about a third is
actually hurting yeah but if you don't
know which ones the third then you don't
know we don't we never know right I
remember like another another Twitter uh
like uh from uh Naval like the famous
Angeles Thunder yeah yeah he developed a
mental model about the complexity of
decisions and his sense is just humans
are very bad at making complex decisions
yeah especially if there's an element of
subjective like you know look I came up
with the idea and I feel so personally
invested in that idea it becomes very
hard like without data yeah to be able
to like you know actually make the right
decision now the with the current
economic climate it's also it becomes
extremely important like you know okay
well which team should I double down on
or which team should I wind down or
which project should I wind down it's
it's important like you know if you just
go and wind down the projects that are
actually beneficial to the product
that's actually worse it never hurts to
like understand the impact of the effort
that you're putting in yeah I feel like
AP testing is the ultimate tool for
achieving intellectual honesty I think
so now
one thing I would say is like you know
thing is the I think the state of the
art the way to identify which ones work
which ones don't now but the problem
with a B testing is that it is a time
consuming process the process of a b
testing comes up with like first you
have to come up with a hypothesis and
then you have to build a variance for
those validating those hypothesis then
you have to ship those variants and then
you have to allocate samples isolate the
experiment run the experiment for like I
don't know two three weeks depending on
like how many samples you have and then
you have to have your data science team
go back and analyze it and then you know
verify it all to like whether you uh
validate or invalidate the hypothesis
that you originally had that entire
process takes so long that most people
don't run experiments as many times as
they should because then what they do is
they save the big decisions for
experiments and the rest of them they
rely on product and tuition what are the
convenient decisions yeah right and so I
think it becomes important to like for
tools to make it so simple yeah it
should be automatic the whole idea of
like every code feature release should
automatically be subscribed into an A B
test and the tools should take the work
or the burden of like analyzing those
and giving you back numbers yeah that
you can then use to make product
decisions
is that the biggest selling point for us
that's it yeah so the the idea behind
statzig is like we think that a b
testing is great but a b testing is such
a time consuming and manual process that
most people don't run as many a b tests
as they should and so it is it is the
important for a tool like static to come
in and say like look you focus on
building features because that's what
you're really good at that's what you
should be spending time on let the tools
take over the idea of like okay any time
we see a rollout and that is that
generates an opportunity for us to go in
and understand okay here's a split in an
otherwise statistically random sample
that we can take the people that are
exposed to the feature as your treatment
and the people that are not exposed to
the feature as control let's compare yep
let's compare all the metrics the
hundreds of metrics and then see if
there's any statistical differences
between those metrics and that is useful
information for you right so that gives
you because why would you build a
feature in the first place you would
build a feature because you believe that
feature is good for your users customers
or business or anything something about
the feature is good and that's your
hypothesis let's validate that so I have
two questions uh I think both are linked
to our Market size or your total total
addressable markets the first question
is uh kind of how many of your current
customers are x Facebook X Google and X
Uber because they have all used EB
testing platform but uh I found it is
difficult for people that never use this
kind of a b testing platform to realize
the value behind it yeah so this is a
really interesting question because
whenever we talk to someone new
um we can quickly tell if we're selling
uh you know in the realm of uh selling
right and are we selling in Tylenol or
are we selling a vitamin I'll explain
what that is like you know
um Tylenol is a painkiller so imagine
someone comes to you and says like look
I have a headache and I need a Tylenol
it's easy for you to like look here's
the Tylenol yeah it's solved for your
pain and I can sell it to you whereas a
vitamin is a little bit different
because you have to first convince
people that it is important for you to
have the vitamin but the moment you try
a vitamin then you feel the benefit of
it and you're never gonna you know not
want it and so
um when we talk to like the the ex
Facebook exuber X Airbnb it is very
clear like they they realize the value
intuitively of a tool like static and it
becomes a much easier conversation and
then when when we talk to uh the other
set of folks there is an element of like
for us we have a lot of awareness to
build we need to like
um be out there do some content uh build
it right you know write blog posts and
go to conferences and talk uh to educate
or to build awareness for a different
way of building measuring and then using
the data to inform your product
decisions and that is a large market and
that we are starting to slowly
um but for us so far the
the success have come from you know the
people that already understand the value
of the tool so we're super early we're
only like 18 months in so right now
we're kind of like in the you know we're
happy just serving the Tylenol Market
yeah eventually we want to address the
Vitamin Market
yeah I think it's uh just uh it is just
a very hard problem to solve like
getting people to realize the value of
vitamins yeah no it is and there are
companies that have done that really
really well I think um you know even
just feature flagging right you know
some of our competitors you know have
done a pretty great job of like
convincing people that feature flagging
is a great way to build products you
know it decouples code shipments code
releases from feature releases you don't
have to tie those two things together
that's a very powerful concept and so
the people before us have done a pretty
good job of like bringing awareness and
now it's on us to like build the
awareness to the next level which is
like it's not just enough if you just
put people you know features behind a
feature flag it's important for you to
understand how each feature is
performing is it beneficial for your
customers your users your business
you talk about our competitors what is
the biggest difference between you and
other like optimizely and I think there
are a couple uh split right yeah so
um like I said one of the things that I
have observed
um was the whenever people talk about
product experimentation the state of the
art is a b testing yeah and that's where
it stops
whereas what we all learned inside
Facebook is like a b testing is is great
but it's it's it it favors Precision
over decision
and so what ends up happening is people
obsess so much about like the Precision
of the A B test and in practice what
ends up happening you need to be making
a lot more product decisions and so
where stats it comes in is like you know
automates the entirety of like running
an A B test so much so that we believe
that when our customers are now running
10 times more experiments than they were
running before I see the efficiency of
running yes experiments Simplicity even
yes so the the biggest
differentiators are how simple our tool
is for you to like just get those
metrics right away two lines of code
right yeah yeah and and then everything
else flows from there the second part is
you don't need to occupy your data
science team to like constantly be
analyzing and making product decisions
the engineers the product managers and
the designers can make these decisions
using static and what it does is like
it's pretty important because it's hard
enough to find good data scientists good
data scientists or see you know there's
not very much it's very difficult and
then what happens is most of the
companies that we talk to hire these
great data scientists and then what do
they put them on they put them on um go
Analyze This A B test yeah and what your
what ends up happening is these these
teams of data scientists are looking
backwards and analyzing and diagnosing
and running queries on an experiment and
then you know making like okay here's a
report and make your product decision
instead those people I mean obviously I
don't think people enjoy doing that they
should be looking forward they should be
like analyzing okay what should the
product be going forward into versus
like looking backwards and so what
static does is also relieves these data
scientists from you know the grunt work
and lets them do creative work which is
what everybody wants to do yeah thank
you for yeah
you know we see so many of our customers
like you you have such an amazing data
science team but what are they doing
they're doing grunt work yeah I think
this is a very good angle like looking
backwards versus looking forward I think
other great data centers wants to you
know help shape the future future
strategy what the product should evolve
into not what the product was already
done like you know so better decisions
yeah looking at the policies only taking
information for the future yeah correct
all right okay the second question also
about total addressable Market we know
at Facebook it's very it's very easy to
do it be testing because of the traffic
size right you have billions of uh
billion of people using the oh yeah but
uh it's not the case for most companies
yes then how yeah this is a common myth
um the myth is that you need to have
large sample sizes like Facebook in
order to run experiments
um and I think that's that's not true at
all so if you're looking for 0.01
improvements in your Mau or BAU metric
or your Revenue metric then you do need
millions of billions of samples the
obviously you know this better than
anyone else which is in order to get
sample size or the um what is the the
minimum yeah you have two factors going
and there's the minimum detectable
effect and then the Baseline conversion
rate as well as the sample size the
turns out that the minimum deductible
effect has a much higher bearing on like
the your statistical power and so when
small companies with like only a few
thousands of samples what they're
looking for is 10 wins yeah 20 wins
they're not looking for point zero one
percent if you're if you're a small
company with only thousand samples if
you're looking for point one percent I
would say like yeah you should stop you
should like go look for larger wins
um and so when you're looking for 10 20
wins you don't need millions of samples
that's true and and so we we we we have
to bust this myth because you know this
comes up a lot of times 39 of our
customers are our B2B companies and B2B
companies they don't have when when they
generally talk to us they're like oh we
don't have that many
um
samples and how do we run experiments
you can definitely run experiments and
these are all like still valid I would
still say like look
you should still put every feature
behind the feature flag and then still
validate the impact of those features
because sometimes what happens is you
introduce a bug Without Really noticing
and those bugs have this map
and those those MDES will manifest
itself as a as a as a change with
statistical significance instantly you
don't have to wait for two weeks it'll
you'll get back that right away and you
should be looking at that and capturing
those and then fixing the the bugs and
not wait until like your experiment is
done to look at those so I I generally
say you know it would be beneficial if
the entirety of the science Community
comes together and it's like actually
busts this myth you need to have Lars
lots of samples yeah I remember a
Facebook I always tell the engineers uh
don't look at the the final metrics
Revenue I mean he's at First Look at
adoption at different layers yeah and
that often tells you more information oh
absolutely that's another good technique
where you look at the top of the funnel
you know correlating metrics you know
things that are that have less inertia
things that can move right away it's
also a pretty good
um you know way to pick the right metric
for how you measure the success of your
own product you know sometimes if you
pick two deep of a funnel metric say if
you pick the bottom of the funnel metric
you know it's very hard to move those
metrics and you don't know if the if the
work you're doing is actually impacting
sometimes people pick like you know 30
day average metrics and then those
metrics have inertia and what happens is
like you know whatever changes you're
making don't immediately manifest you
have to wait alongside
so my general philosophy is like you
know pick top of the funnel metrics
generally the ones that are correlated
to the ones that you want to move and if
they're moving in the right direction
then that's a good sign
foreign