Starting a Career in Data Science in the Age of AI

Starting a Career in Data Science in the Age of AI

人工智能时代:如何开启数据科学职业生涯

I was recently asked by a college student in data science and computer science for my advice on entering and succeeding in the field of data science and machine learning today, without selling your soul or sacrificing your ethics. It was a really hard question to answer, because I myself entered the workforce almost 20 years ago, and the field of data science 10 years ago, and so much has changed since that time. Questions of ethics were much less pronounced (although not absent) when I became a data scientist, and today they are front and center. I did my best to answer in the moment, but I’ve spent some more time thinking about it, and I feel like there are a few key considerations for students trying to map a path to a rewarding career in our field that will last.

最近,一位数据科学与计算机科学专业的大学生向我请教,询问如何在当今的数据科学和机器学习领域入行并取得成功,同时又不违背良心或牺牲职业道德。这是一个非常难以回答的问题,因为我本人是在近 20 年前进入职场,并在 10 年前进入数据科学领域的,而自那时起,这个领域已经发生了翻天覆地的变化。当我成为数据科学家时,伦理问题虽然存在,但远没有今天这样突出,而如今它们已成为核心议题。我当时尽力做了回答,但事后我又深入思考了一番,觉得对于那些试图在这一领域规划长期且有意义的职业生涯的学生来说,有几个关键点值得考虑。

Diversify Industries

多元化行业选择

First, working in the software industry is not necessarily the path for all of us. Data science broadly, and machine learning engineering in particular, have fantastic applications in pretty much every sector, much more so than in the past, so I encourage students to consider fields like healthcare, government, nonprofits, and other sectors. Data can make a massive impact on the success and efficiency of all kinds of work when analyzed well and modeled in a sophisticated way. Internships are a good way to find out what data scientists in different sectors and industries do, and I encourage students to just simply read the job descriptions as well, and talk to practitioners when possible, to help get a feel for the expectations. The character of the roles will change (often rapidly) but there’s no better way to figure out what the job can be like than to ask people who are doing it.

首先,在软件行业工作并不一定是所有人的必经之路。广义上的数据科学,尤其是机器学习工程,在几乎每个行业都有极好的应用前景,这比过去要广泛得多。因此,我鼓励学生考虑医疗保健、政府、非营利组织等领域。当数据得到良好的分析和复杂的建模时,它能对各类工作的成功和效率产生巨大的影响。实习是了解不同行业数据科学家工作内容的有效途径,我也鼓励学生多阅读职位描述,并尽可能与从业者交流,以了解行业预期。职位的性质会(经常是迅速地)发生变化,但要了解一份工作究竟是什么样,最好的办法莫过于去问那些正在从事这份工作的人。

Data Education

数据教育

On the topic of ethics in the practice of data science, I believe it’s the responsibility of data scientists to educate colleagues about ethical, effective data utilization. I’ve written about this many times, but we are the people in prime position to help whole organizations understand how to use data and models, including LLMs, in ways that produce results and don’t endanger people’s safety, health, privacy, or welfare. It’s incredibly, horrifyingly easy for organizations without proper data expertise to careen ahead with AI and other technologies with ignorance of the risks, and it should be part of the practitioner’s role to guard against that, just like it’s an accountant’s job to prevent the company from committing financial crimes. If you want to feel good about your work and sleep well at night, taking this responsibility seriously and being a voice for ethical, safe applications of machine learning is a good place to start.

关于数据科学实践中的伦理问题,我认为数据科学家有责任教育同事如何合乎道德且有效地利用数据。我曾多次撰文探讨这一点,我们处于最有利的位置,可以帮助整个组织理解如何以既能产生效益,又不危及人们安全、健康、隐私或福祉的方式使用数据和模型(包括大语言模型)。对于缺乏专业数据知识的组织来说,在忽视风险的情况下盲目推进人工智能和其他技术是极其容易且可怕的。防范这种风险应当成为从业者职责的一部分,就像会计师有责任防止公司犯下金融犯罪一样。如果你想对自己的工作感到心安理得,并睡个好觉,那么认真对待这一责任,并成为机器学习伦理与安全应用的倡导者,是一个很好的起点。

Solving Problems

解决问题

What else should you expect to do in a job as a data scientist or MLE? Well, expect to be solving problems. That’s really where data science shines, where business problems or uncertainties appear, and knowing the answer is vital to making a correct decision. You won’t start out finding your own problems, instead your manager and leadership will be defining and scoping the questions that need answering, but in a good role, you’ll have some autonomy in choosing the strategy and methods you use to answer. By doing this, you’ll learn how to approach different kinds of data questions, and you’ll see how the problems are defined and where they come from, and this kind of experience is invaluable. That’s what differentiates entry level from senior level practitioners in data — the ability to recognize a problem’s archetype, hunt through messy and complicated data to develop a plan, and then to execute on it, generating a statistically rigorous and sound answer. This skill set, in my experience, can take you to most any industry or sector that interests you, because the real skill is problem solving. You’ll learn a lot in college (if you do it right) and you’ll need that knowledge to progress, but you can’t learn the real problem solving skills until you’re out there in the field doing it, in my experience.

作为数据科学家或机器学习工程师,你还应该期待做什么?答案是:解决问题。这正是数据科学的闪光点所在——当业务问题或不确定性出现时,找到答案对于做出正确决策至关重要。起初你可能不会自己寻找问题,而是由经理和领导来定义和界定需要回答的问题,但在一个好的岗位上,你在选择解决策略和方法时会拥有一定的自主权。通过这种方式,你将学会如何处理不同类型的数据问题,并了解问题是如何被定义以及从何而来的,这种经验非常宝贵。这就是数据领域初级从业者与资深从业者的区别所在——识别问题原型、在杂乱复杂的数据中寻找线索以制定计划,并执行计划,从而得出统计学上严谨且合理的答案。根据我的经验,这套技能可以让你进入任何你感兴趣的行业,因为真正的核心技能是解决问题。在大学里你会学到很多(如果方法得当),这些知识是你进步的基础,但根据我的经验,真正的解决问题能力只有在实战中才能习得。

Don’t Over-Specialize

不要过度专业化

In contrast, training and building frontier models is not the way to a lasting career for most. Most of us working in the field will never get near “cutting-edge” frontier model construction, and I for one wouldn’t want to. Technologies change, and methods of performing machine learning tasks change — I have witnessed this firsthand. Some narrow specialists will focus on training LLMs, but I recommend flexibility and versatility, especially in your early career, so that you have more options as the field inevitably changes around you. I am glad that I have a good foundational knowledge of neural networks and can train and tune them, but I am equally glad that I also know GBMs, NLP, clustering, and other techniques, because these are just all tools in my toolbox, which I can use when they’re appropriate. LLMs are not going to be the end of machine learning, and we need to be watching for whatever comes next.

相比之下,对于大多数人来说,训练和构建前沿模型并不是实现长期职业发展的途径。我们大多数从事该领域工作的人永远不会接触到“尖端”前沿模型的构建,我个人也不想这样做。技术在变,执行机器学习任务的方法也在变——我对此有亲身体会。一些狭窄领域的专家会专注于训练大语言模型,但我建议在职业生涯早期保持灵活性和多面性,这样当领域不可避免地发生变化时,你会有更多的选择。我很高兴自己拥有扎实的神经网络基础知识,能够对其进行训练和调优,但我同样庆幸自己也掌握 GBM(梯度提升机)、NLP(自然语言处理)、聚类及其他技术,因为这些都是我工具箱里的工具,可以在适当的时候使用。大语言模型不会是机器学习的终点,我们需要时刻关注接下来会出现什么。

Use LLMs Cautiously

谨慎使用大语言模型

You might also be wondering how and when you should expect to be using LLMs in your professional life, and this is a question most everyone in white collar fields is asking, not just data scientists. I struggle with this, because AI is a useful tool in many situations — writing code, for example — but abdicating our critical thinking processes to the chat bot is dangerous. I’ve talked about this elsewhere, including discussing the financial implications, and I encourage entry level folks, at least for now, to make sure they have the capability do the job manually, even if they don’t have to every day. I don’t recommend taking a rigid stance of refusing to use it, because in coding at least you really may fall behind peers, but I propose finding ways to use it for what it’s good for, and being selective. I routinely get complimented at work because I write all my own text (like these articles, where I never, ever let AI touch them) because it sounds and feels human. People appreciate the fact that I care enough about the work and about my audience.

你可能也在思考在职业生涯中应该如何以及何时使用大语言模型,这不仅是数据科学家,也是大多数白领领域从业者都在问的问题。我对此也感到纠结,因为人工智能在许多情况下确实是有用的工具(例如编写代码),但将我们的批判性思维过程完全交给聊天机器人是危险的。我在其他地方讨论过这个问题,包括其财务影响。我鼓励初级从业者,至少在目前,要确保自己具备手动完成工作的能力,即使不必每天都这样做。我不建议采取拒绝使用它的僵化立场,因为至少在编程方面,你确实可能会落后于同行,但我建议找到它擅长的用途并有选择地使用。我在工作中经常受到称赞,因为我所有的文字都是自己写的(比如这些文章,我从不让 AI 触碰),因为它们听起来和感觉上都更具人性。人们欣赏我足够重视工作和受众这一事实。