课代表立正的官方Podcast
深度访谈,有用干货,亲身验证的「真本事」
Superlinear Academy创始人,Maven Top AI Instructor
前Statsig布道师(OpenAI收购),腾讯副总监,Meta,Amazon;康奈尔经济学博士
社区:Superlinear.Academy
课程:ai-builders.com
个人:lizheng.ai
hi uh
I don't remember my English opening
hi folks welcome back to my Channel
today I'm aesthetic interviewing Tim Tim
is the actual team why don't you give a
high idle data science at statsig I've
been with the company from the very
start in February of 2021. so prior to
this I was a data scientist at Facebook
so where I spent uh I think like five
total five years working on Facebook
core app in in gaming and also I worked
a little bit in Facebook reality Labs so
doing some of the more Cutting Edge
metaverse style projects so I tell
people that I've worked on
products that I've had as as large as a
billion users and as small as just like
10 000 users which one is that would be
like the Facebook reality Labs yeah or
they are having that many users today
what is the one billion uh so that would
be like on Facebook gaming it's just
like you had that many eyeballs on
gaming content it was our opportunity to
sort of capture that and drive
engagement I was uh working at the
startup yeah
um it's been it's been a lot of fun it's
been pretty intense I think I've got to
try I've got to I've been able to learn
a bunch of whole a bunch of new skills
that I would have never had the
opportunity before so uh being the only
data person on the team
um the only at the beginning I was the I
was the only data person on the team is
like um I had to be more than a data
scientist I had to be also a data
engineer I also had to be a data
architect and designed that I also had
to be
um thinking there was no one else
thinking at the time about things like
marketing or or business positioning or
our strategy like even those sort of
things sort of I was able to contribute
to
um
at the early start
data people to have we now we have uh
three data scientists and we have a lot
of Engineers working on the data
infrastructure side
how to define Engineers working on data
infrastructure or resistance engineer
yeah it's it's actually a blurry line
yeah um I think it's I don't think
there's actually a really clear
definition
um and I think there's a
um that there's a lot of shared
responsibilities between the two
pleasure okay so early on you were
basically uh
Jackal file Trace uh and uh uh but
middle life as a data scientist what is
a unique uh value or unique contribution
that no one else can uh replace
at static aesthetic that's static yeah I
think um the way I view is is that
because we're selling an Analytics tool
um we not only are there to engage with
other data scientists
um because other data scientists want us
to be able to speak with somebody who
understands data science and this is
partly they have to believe that we've
set up the analytics and the
computations correctly yeah and so we're
the ones that we we do that and we're
also the ones that are able to
communicate that so that's like the
first thing
um the second thing is like we are also
the we also represent the Viewpoint of
the customers when building our tool
both from there's two perspectives one
is like what would a data scientist want
to see in the product but also data
scientists are a little bit like almost
guardians of like data at their
companies and how people use data and so
they have to be able to trust that we
are putting a tool in the hands of PMs
and engineers and they and data
scientists want to make sure we're doing
it
Faithfully in a way that doesn't cause
them more work
because there's ways to misinterpret
data that could actually cause like
people to have wrong insights and then
in that point like data scientists are
now struggling to actually squash
problems that we've created yeah instead
of us solving problems for them
right they're so like I think there are
two main follow-up questions uh from
this was statistics Mission or Statics
uh differentiator with other AP testing
companies like uh split split split wise
or whatever optimize
other competitors is you want to make a
b testing easier uh that's less
complicated or less academically right
uh
I feel like that creates some tension
between uh like giving this to two
engineers and they can just run it
versus a data scientists want to have a
rigorous process and rigorous thinking
uh after
um like with the a b testing the how do
you balance that and the second question
is uh uh in order to convince other data
scientists or your customers data
scientists that we are providing the
right tool you basically have to know
more than them right do you need to feel
that to be a necessity but I think I'll
start with like your your first question
over like how do we balance making
experimentation accessible but also
being like rigorous yeah and that is
that is our the biggest challenge we
have as a data science team actually I
think that's the biggest challenge we
have as a company in general
um because our tool is fairly
sophisticated from a stats and analytics
perspective
and yet we're trying to put this
powerful tool in the hands of people who
may not fully understand statistics or
maybe not know the definition of a
p-value I think a lot of what we're
building is fueled by the fact that
we've seen this work at Facebook or meta
that like you can put a powerful
statistical tool in the hands of people
who aren't data scientists and and have
a really power product development and
so we've tried to capture
some of that nicety but also adding our
own flavor on top of that it is tricky
and it is something that like we we
actively think about I what we try very
hard to is our principle is that
um
we
we will give you a lot of flexibility
but we will try to guide you through UI
and design to best practices so things
like making sure that uh when you are
deciding how long to run an experiment
for that there's a power calculator
that's readily accessible things like if
you're going to make an early decision
we actually let you know that like hey
it's possible this experiment is not
fully powered yet just little nudges
like that into the UI and design that
sort of like guide people where you may
not fully understand like the peaking
problem and like the nuances of it but
the fact that our UI is sort of like
sort of like pushing you to run this
experiment a little longer than you may
be or that just to let you know a
warning that you are running making an
early decision these are just kind of
the things that we've tried to build in
from an experience what kind of
knowledge or understanding does it take
for someone to use it to correct
somebody we saw the UI Nations right
like not knowing the definition of
p-value or not understanding the true uh
implications of p-value like is it uh
okay to yeah I think this is we
discussed this heavily uh a lot um
because like ultimately we are showing
frequent test results that are that have
P values and confidence intervals and we
know what those mean to statisticians
and that might not necessarily be the
takeaway that someone who doesn't
understand type 1 and type 2 errors will
take away but what we do want to give
let people know is like for example like
subtle indicators like making a result
green signify that this is actually the
result you are looking for that this is
like you know we all know this means
it's a low P value underneath your 95
significance threshold and while while
an experimentalist may not actually know
that full definition
seeing a green result should give you
some confidence that that is actually
what you're trying to look for when
you're running this we also do things
like making sure that we encourage folks
to list out their key metrics up front
that's experimental best prep how much
is the information lost
yeah so we we do this um the way our UI
is designed and probably this is best
for a lot of like really technical
products is that we actually show you a
very simple simplified version of your
data where it's very distilled down and
there's like very visually you can sort
of like instantly get some very key
takeaways but what we hope and and this
is especially true in experimentation
it's a lot of like peeling an onion
where like whenever you see initial
results
at least data people will have like
follow-up questions immediately yeah and
you'll want to drill in on those and so
we actually provide in the UI like very
easy ways to go deeper and deeper
um into more complex things into more
complex statistics so we don't we
haven't dumbed down our product like we
we've simplified the initial version but
if you wanted all the complex stats and
the drill down such as being able to see
time series or filters by by cohorts
like that's all available okay um in
there um and readily accessible
um that is actually a good point because
I feel like I I say it's a Facebook or
other companies that work at uh when you
when people see uh experimental results
data scientists back to a roadmap uh
post as well did I scientists you want
to have hypothesis they want to
understand why this result happened
right versus a lot of PMs and Engineers
they're looking for specific results
they're looking for the grain yeah they
got the Grain and good to go if they
don't get the green they want to
whatever I was gonna try to make it
great and sometimes they don't have
hypothesis behind it
yeah this is one area that we really
want to train our customers it's like
you don't want to just look for the
green you want to be able I I call it um
telling a complete story that like I
mean that you're that whatever results
you see should be consistent with the
change that you made and it should be
somewhat expected if you have something
very unexpected you should question that
result
um and so we try to make sure that
that's why we ask our folks to
put their hypothesis down what they
expect to see and quote some metrics
that they expect to move things should
be somewhat consistent you shouldn't
just look for green things
um just to make that decision yeah I
also feel like when people have
hypothesis yearly or estimate uh how how
often they are right like absolutely
which is that one third but uh in the
online control experiments like
mountains of the hypothesis are as
expected or most of the changes they
either are insignificant or even like
are proven to be negative so people
would have frustration with it when they
don't see the Grain and sometimes the
frustration can be a question about the
tools like I was right and your two is
not showing me green so your two must be
wrong
yeah we have this happen all the time I
find that experimentation is like super
humbling not only in like the number of
ideas that you could have that work but
also in the magnitude like very often
times like if somebody's never done
experimentation they'll say like oh yeah
this change is so obvious this is gonna
be so good it's gonna be plus ten
percent to revenue and then they run it
and it may still be good but it's like
plus one percent and it's just like you
have to I think like everyone who does
experimentation sort of like knows how
to calibrate downward
um it never calibrates upward um but we
also find that uh some companies also
just have a higher batting average so
usually small companies that are working
on unoptimized products usually there's
a lot more wins around but whereas if
you're at a big company like Facebook
level
um the number of good ideas that are out
there is far fewer so yeah so it's a lot
harder to um to find those wins but yeah
it's something we have to temper
expectations of our customers
um that like hey this is like the fact
you are seeing a negative result or a
neutral result
um is not a bad thing like you actually
now are able to measure your impact it's
up to you what you do with it
um I feel like this would be way easier
if the like your customer have seen uh
like work at Facebook for example like
I've seen this work and it would be way
harder if they never run experiments
before or like crack experiments before
yeah uh like the latter problem is it
even sellable like you're 100 correct I
think our biggest challenge is
um
getting people to realize the value of
experimentation I think if people have
worked at a company like Facebook or
Google or Microsoft they sort of
understand the value of experimentation
um but if you're at a but if you've
never been exposed to that and you read
things online I think it's harder to
convince you
um that this is a good thing yeah but
that's our challenge is like to tell
show people how it's done at big
companies medium-sized and small
companies and what sort of learnings and
how people are using it today and why
and how that's driving product
development
have you had examples of people not
believing interface or very skeptical at
first and then over time started to
trust anymore
we're starting to chip away at that
we've have I don't think we've had
somebody who's completely skeptical use
our product and be convinced the other
way but we've had people who have been
um listen we want to try statsig for
just rolling out features and just the
fact they're seeing measurement and data
now they're realizing oh we can actually
start to make decisions on this that
they may not have in the past and so I
think just by the way I view it is like
we're just shining a flashlight on like
what your features are actually doing
it's up to you to take a look it's up to
you to interpret it and it's up to you
to make decisions but I think just
showing people The Shining Light on
these on this data and results has been
pretty powerful so far in learn getting
people to realize the value it is a
necessary First Step but uh like
unleashing the power or the full
potential right I feel like there are
several steps for example you need to
just measure which which one is positive
and which one is seductive but then you
need to hypothesis and understand a
product and guide it to a better
Direction that's the Forward Thinking
part comes in I think we already
transitioned into the second question so
let's just just go there what how do you
feel uh the I guess Talent density well
whatever like in order to do a special
step two right like realizing this uh
the realizing the power of experiments I
feel like this power is step one and uh
unleashing the power by having
hypothesis and having the experiments
not only to tell you which one is
positive but also to uh help you learn
your customer better or learn your
products better I feel like this that
takes talents right
um
yeah I think that the experimentation
journey is actually like fairly long
like when you first start using it you
can operate at one level but it but you
have to become sort of like practice and
hopefully like having a reliable tool
supports that but there's definitely
many light layers
um you can get to an experimentation to
the point where you're actually
comfortable just trying random ideas
yeah um and being able to just build it
quickly and test quickly
um I think that to me is where
was where what we saw happen at Facebook
and what we think companies can get to
eventually yeah but I think you're right
that people have to take small steps um
in order to get to that point yeah my
question is about do you see the Second
Step being done at uh companies or other
other than this big tech companies like
having hypothesis around the experiments
interpreting the results to understand
the product and the people the customers
instead of just you know this is
positive that is next let's move out
yeah I think when we guide people that
I'm there we actually tell people that
you should be focusing on metrics that
matter to you you should be focusing on
the user experience and the metrics that
you see go positive negative or neutral
should be consistent with what you think
the user is experiencing
um I think we different companies are at
different stages I find that if there is
a data scientist on the other side they
can actually use statsig to guide their
teams to how to interpret this and I
think that's usually a faster process
but then we also have found that like
purely engineering teams are able to
sort of
view metrics
with a degree of skepticism and so
therefore start actually being able to
interpret it properly but that's a
harder I think it's easier if you have
some somebody to guide that and that's
partly the role of a data scientist all
right so the question is about the
distribution of data centers at a
different stage of companies right you
said you have a spectral mixture of
different companies so very early uh I
guess serious B the pre-ipo and the big
appeal companies I don't know if that's
the right kind of stages but the
distribution of data centers yeah I I
sort of linked this to what I call data
maturity and it's very different it's um
there's not really a set formula that
like once you hit 20 Engineers you must
have your first data scientist we've
seen some companies bring on a data
person very early on and we've seen
other companies like not have any data
people so they're getting to be like
fairly large there's not really a set
formula
I think for me and and we have to work
with all companies all across the
Spectrum I think like and so the way I
see like sort of the data maturity going
is that you usually you'll get your
first data person that will actually
start putting helping your
you know the real value of a data
scientist to me is like
helping the rest of your team up level
their analytics ability helping them
understand how to interpret data like
the data scientists themselves can
produce insights but I think the real
power is actually educating your team
how to use data properly data maturity
so like uh at the different stages of
data maturity to the early first hire a
data engineer
I think so if you go by the book like
I've seen like best practices that
usually you bring in a data engineer
first so they can sort of at least lay
down the groundwork that like here's a
few data sets and
um engineers and PMs can sort of like do
some basic queries on there in order to
be able to get some value out and so the
usual data scientist comes after
um I have seen times where your first
data scientist has hired before a data
engineer but their work
that they started doing at the beginning
is data engineering work yeah
I heard this happen a lot at
pharmaceutical companies actually like
it's actually respond to do uh deep
learning they want to do AI so they hire
a lot of AI scientists and an app doing
detention networking yeah I think it's
one of these things where if you don't
have that function at your company it's
hard to that first person you hire is
actually hard to know what you're
looking for it's actually how does a
person interview
someone in another field and know that
they're going to be a fit for the
company it's like it's a little bit
tricky yeah yeah so what is uh
advice for data scientists who want to
uh try working on to start up because
it's very easy for engineers to jump
from a big Tech to startups but uh
perceived that it's very hard for data
centers to do so like for example data
centers that big companies they are used
to this nice infrastructure clean data
and they start to yeah I think it's I
saw I would have two um pieces of advice
and one is just have a very much a
startup mentality where you are prepared
to roll up your sleeves and I tell
people like get dirty with the data like
there is no data task that is built
beneath you
um if you're at a startup so like even
just cleaning data removing nulls like
just these the basic grunt work like
there's no one else to do it you have to
do it
um that is like the first thing is just
like let's start the second one is just
and maybe it's similar but be prepared
to like learn new skills be willing to
learn new skills be be willing to try
things that are outside your comfort
zone and so like for example one of the
things we had to do was learn how to set
up a how to set up spark as our for our
compute and that is not something which
I would have ever thought I would know
how to do uh benefactor back at Amazon I
spent a week setting it up a setup
sparked query data from other data from
other words okay nice but uh like I know
that has to take a data engineer at most
half a day
yeah yeah
and it's but it's it's one of these
things where like there is like and I
fully know that somebody else could have
done this job faster but there's nobody
else around yeah you are the person so
like you have no choice but to roll up
your sleeves and just find a way to do
it and learn and find the quickest way
to do it like you know you haven't done
it well but it's good enough to move
forward so as uh it doesn't hurt to be
more footstep especially if like for
startup you have to be food stamp kind
of I yeah I think like so when I joined
statsig I was not full stack but I think
so I think my recommendation is just a
willingness to learn
um or just in a
can-do attitude like to know like yeah
yeah
that's exactly it yeah so just like find
a way to do it
um and even if it's outside your comfort
zone yep yeah
because you deal with a lot of startups
who needs data scientists basically
right at least they need the tune to run
experiments or do you think it's going
to be a viable path or uh like growing
path for data centers or to join startup
companies
or do they feel like yeah two can just
uh you know satisfy most of them
um I yeah I think uh it depends what it
is that your company is looking for but
I find the way I phrase what static does
is we actually up level your data
scientists
um when when I was at Facebook and maybe
you had the same experience uh I and I I
had interviewed like hundreds of other
data scientists and I would always just
as a warm-up question just ask them like
hey so what how's your current company
and why are you looking for a job and
more than half the times they would tell
me like hey I'm like the only data
scientist at a small startup and all I'm
doing is crunching confidence intervals
yeah and and and and it sucks and I've
been there
um I know like it's super tedious it's
one of those things where like if you
mess up one small part of the
calculation you'll get the opposite
insight and so you have to like triple
check quadruple check these things it's
like it's not really it's not fun and it
doesn't feel like
a company is getting their full value of
their data scientists if they have them
just crunching confidence interval so
data scientists don't want to do that
job and the company probably doesn't
want their data scientists doing that
job so I think our tool can come in and
sort of take that very Elementary level
of analytics burden off of the data
scientists so now the data scientist is
not crunching results they're now
focusing on more important problems such
as designing the experiment interpreting
the experiment
um looking for follow-ups like diving
deep into the data and seeing what else
other insights there are I think like
those that's where a data scientist
starts bringing a lot more value to
their company than just crunching a B
test results
so it goes up by making experiments
easier and easier to run right and
unleashing the power of data centers
that actually have or understand this
contributing more to this kind of these
companies and hopefully we can have more
data scientists working at these
companies as well yeah yeah all right
that's it that's it okay thank you Tim
awesome to have any other things you
want to have uh no thank you very much
also heard static is uh hiring to the
center ah yes what kind of uh we're
we're hiring data scientists and data
Engineers uh and I think as long as we
continue to succeed we will continue to
be hiring because even if we don't have
actually job openings but the data
scientists we're looking for this sort
of like
two skills and
they and one is like being able to work
with customers directly so communication
skills being able to translate
statistics or ex complicated
experimental Concepts in Easy plain
English is is a pretty important one the
second one is uh more on the technical
side so like being able to put ideas and
and data science Concepts into code so
being able to write in SQL being able to
write in spark and so that might be I
would call a little bit more of like a
statistical engineer or an analytics
engineer yeah but those are the sort of
two skills they're kind of on the
opposite side of the spectrum a cloud
statistic maybe applied statistics yeah
essentially like we need people who work
and work in the code base and ship
production code nice yeah okay also
about data Engineers
um data engineering is assembly do you
have any we don't have any we're looking
um yeah I think like data modeling is
something we need uh to bring in so you
ask what the difference is between an
engineer that does uh works with data
and a data engineer and I mentioned
there's a lot of overlap I think one
thing that doesn't overlap well is
actually data modeling and I think like
data Engineers who are really good at
data modeling is something we're looking
for all right okay thank you if you are
interested uh you can find the team uh
linking yeah absolutely please all right
thank you thank you bye