生态认知困境:信息缺失与行动迟缓
想象一下,如果你是一位医生,却只能看到病人身体的五分之一,你将如何诊断并进行治疗?这正是我们目前在面对地球自然生态系统时所处的困境。我们迫切需要采取行动保护受威胁的生态系统,但对于地球上的生命,我们仍然知之甚少。生态学家和AI研究者Sara Beery认为,AI的潜力在于能够指数级地提升我们对物种和生态系统的认知。然而,要实现这一目标,我们需要改变AI在生态学中的应用方式,转向更加灵活、交互性强,且能够帮助科学家从数据中发现隐藏知识的方法。
Original English
Imagine you're a doctor and you're trying to save the life of a patient, but you can only see a fifth of their body. How are you going to prescribe medicine? How are you going to do surgery? See, this is exactly the situation we're in with nature across the planet. We need to act now to protect ecosystems under threat, but there's so much we don't know about life on Earth. I'm an AI researcher and an ecologist, and as a professor at MIT, I lead a research group that develops methods to help us learn more about the natural world. And I see a future where AI can help exponentially increase our ecological knowledge across species and ecosystems. But to get there, we need to change how we use AI in ecology. We need methods that are flexible, methods that are interactive, methods that scientists can use to discover knowledge hidden in our data.物种多样性危机:未知的八成
科学家估计,地球上共有1000万个物种,但我们目前只观察到其中的200万个。这意味着地球上80%的生物多样性仍然是未知的。仅仅知道一个物种存在是不够的,我们需要了解它的栖息地、食性、迁徙模式等更深层次的信息,才能有效保护它。例如,如果北美洲的昆虫数量骤减,这对以昆虫为食的鸟类意味着什么?哪些鸟类将面临最大的风险,哪些鸟类能够适应其他食物来源?而更高级别的捕食者又将受到怎样的影响?生态系统中的每一个物种都是相互关联的,对一个物种或一组物种的威胁,可能会引发整个生态系统的崩溃。
Original English
Now let me tell you why this is so important. Scientists estimate there are 10 million species sharing the planet with us. But we have only ever observed two million of those. That means eight million species, 80 percent of the diversity of life on Earth, remains unknown. And we need to know much more than just a species exists to be able to protect it. Where does it live? What does it eat? Does it migrate? How far? This deeper knowledge about species takes far more than a single observation. But it's necessary to understand what puts species at risk. So, for an example, what if insect populations crash across North America? We know this is currently happening. What does that mean for birds that eat insects? Which birds are going to be most at risk, and which are going to be able to adapt to other food sources? What about predators further up the food chain that eat birds? Everything is interconnected, and a threat to one species, or a group of species, can ripple outward and trigger the complete collapse of an ecosystem as we know it.多重威胁下的物种加速灭绝
不幸的是,物种正面临着来自四面八方的威胁:栖息地缩小、气温升高、食物和水源减少、自然灾害频发,以及外来物种的入侵和竞争。因此,目前的灭绝速度是基于过去数据的100到1000倍。全球的科学家、政策制定者和社区成员正在努力了解导致这一现象的原因,以及可以采取哪些行动来阻止它。然而,我们常常是在发现物种的同时,也在记录它们的灭绝。例如,2017年才被发现的塔帕努利猩猩,是地球上仅有的三种猩猩之一,但它在被发现之前就已经濒临灭绝。
Original English
Unfortunately, species are under threat from every direction. Habitats are shrinking, temperatures are rising, food and water sources are disappearing. Natural disasters like fire are causing large-scale death and displacement. And invasive species are moving in and outcompeting native species for resources. As a result, extinction rates are now 100 to 1,000 times higher than what we would expect based on past data. Scientists, policymakers and community members worldwide are racing to understand what is causing this, what are the factors that are most contributing to this loss and what actions we can take to stop it. But unfortunately, it can feel like we're discovering species just in time to write their obituaries. Take the Tapanuli orangutan. We discovered this orangutan in 2017. It's one of only three species of orangutan on Earth, and it was critically endangered before we even knew it existed.海量生态数据的潜力:iNaturalist的案例
传统的生态数据收集方式过于缓慢,无法应对当前的危机。幸运的是,我们已经拥有了大量的生态知识数据库,但我们对其的利用程度还很有限。以iNaturalist为例,这个平台已经上传了3亿张图片,由热情的志愿者识别了其中的物种。这种物种出现的数据本身就对科学产生了变革性的影响。然而,在这些像素中还隐藏着大量的知识宝藏。例如,一张被标记为格兰特斑马的图片,不仅证明了格兰特斑马在该地点和时间被观察到,还显示了三只格兰特斑马,并且可以根据其独特的条纹图案识别每一只个体。通过识别个体,我们可以监测物种的迁徙、研究物种的社会网络、评估其生长和健康状况,甚至估算整个种群数量。
Original English
Traditional forms of data collection are just too slow to keep up with our current crisis. And this is where I finally have some good news, because we are sitting on vast databases of ecological knowledge, and we have barely scratched the surface. Let's talk about just one of these databases, which is a platform called iNaturalist. 300 million images have been uploaded to this platform by passionate volunteers. In every single image, the community has identified a species, and that level of species occurrence data has already been transformative for science. But there is a hidden treasure trove of knowledge that remains in the pixels. So let's look at just one of these images. In iNaturalist, this was labeled Grant’s zebra. And it's clearly evidence that Grant's zebra were sighted in this place and time. But it shows us so much more than that. There are three Grant's zebra in this image. We can identify each of them to the individual level based on their unique stripe pattern. By identifying individuals, we can do things like monitoring how species move across the planet, looking at social networks of species, growth, health, even estimating the full population size.AI赋能的生态知识发现:Inquire系统的突破
这些斑马还与一头角马共存,甚至可以看到一只牛椋鸟,这种鸟类以蜱虫为食,有助于减少疾病的传播。我们还可以观察图像背景,识别植被的类型和覆盖率,估算生物量,并了解当地储存的碳。通过分析一张图片,我们就能获得如此多的知识,将其乘以iNaturalist中的3亿张图片,再加上其他生态数据库,如xeno-canto中的数百万条生物声学记录、wildlife insights中的数千万张相机陷阱图像,以及FathomNet中的数千小时深海录像,我们正坐拥一座生态金矿,而提高获取知识效率是关键。如果假设每查看一张图片需要一秒钟,那么仅仅查看iNaturalist中的所有图片就需要全职工作40年。人工智能的出现,正是为了帮助我们快速浏览这些数据。
Original English
These zebra are also coexisting with a herd of wildebeest. And if we look closely, we can even see an oxpecker, a bird that eats ticks and helps reduce the spread of disease. We could look at the background of the image and identify the type and coverage of vegetation. We can estimate biomass, use that to learn about locally stored carbon. We can look at what the animals are eating in the image and build a stronger knowledge of a local food chain. Take this much knowledge in one image and multiply it by 300 million images in iNaturalist, and then add in our other ecological databases. Millions of bioacoustic recordings on xeno-canto, tens of millions of camera-trap images and wildlife insights, thousands of hours of deep-sea footage in FathomNet. We're sitting on an ecological goldmine and the problem is accessing the knowledge efficiently.Inquire:无需训练即可直接提问的AI系统
例如,一位生态学家对鸟类的食性感兴趣,希望找到数据库中鸟类吃昆虫的例子。他们可以训练一个AI模型来帮助他们,为此收集数百甚至数千个例子来教模型识别目标。然而,每次需要寻找新的信息时,都需要收集大量的例子,这仍然太慢。因此,我们需要重新思考问题:科学发现始于科学的好奇心,在于对世界如何运作的提问。例如,格兰特斑马可以迁徙多远?火灾后哪些植物会重新生长?鸟类在冬季吃昆虫吗?如果我们可以直接向数据库提问并获得答案,那将是多么美好。为了实现这一目标,MIT的研究团队开发了一个名为Inquire的系统,它能够帮助生态学家在无需收集任何训练样本或编写任何代码的情况下,从数据中找到答案。
Original English
So say you want to look through all this data, assuming it takes you about a second to look at every image, you would need to work full-time for 40 years to look through all the images in iNaturalist alone. And this is where AI is transformative. It can just help us look through all the data quickly. So an ecologist today, say, they're interested in bird diets, and they want to find examples of birds eating insects in the database. What they can do is they can train an AI model to help them. So to do this, they collect hundreds or even thousands of examples to teach the model what to look for. Now once they’ve trained this model, it’s an incredible tool. It can very, very quickly find new examples of birds eating insects in the database. But this process of collecting hundreds or thousands of examples every time we want to look for something new, it's still too slow. So let's reframe the question. Scientific discovery really begins with scientific curiosity, with asking questions about the world and how it works. Things like, how far can a Grant's zebra migrate? What plants grow back after a forest fire? Do birds eat insects during the winter? Wouldn't it be great if instead we could just directly ask questions to our databases and get answers back? This is what my team at MIT has been working towards, and we've developed a system that we call Inquire that helps ecologists find answers in the data without collecting any examples to teach an AI model or needing to write any lines of code.Inquire的工作原理与应用前景
Inquire系统的工作原理是开发能够学习和理解图像与科学语言之间相似性的AI模型,从而实现直接提问。生态学家首先设计实验,将科学问题分解为一系列搜索词,例如“鸟吃昆虫”。Inquire系统会将这些搜索词与3亿张图片进行比较,并在几秒钟内完成。该系统经过优化,能够快速高效地进行搜索,并且所需的计算能力远低于ChatGPT等生成式AI方法。搜索完成后,系统会根据相关性对图像进行排序,方便科学家专注于最可能相关的图像,并快速验证匹配结果。科学家可以导出经过人工验证的数据进行分析。一位合作者使用该系统发现了数千个鸟类吃昆虫的例子,以及种子、水果、坚果、腐肉、花蜜、植物等。他们分析了夏季和冬季物种食性的差异,发现一些鸟类在冬季也吃昆虫,例如美洲知更鸟,但数量远少于夏季。而另一些物种,如美洲树麻雀,在夏季高度依赖昆虫作为食物来源,但在冬季则完全不吃昆虫。整个过程,从提出问题到获得答案,仅用了三个小时,而另一支团队手动整理数据进行类似研究则花费了1560小时。
Original English
Now under the hood, what we're doing is we're developing AI models that can learn and understand similarities between images and scientific language. And this is what allows us to just ask. So how does Inquire work? Well, first, an ecologist designs an experiment by taking a scientific question and breaking it down into a series of search terms that they can use to discover data that they'll analyze downstream. So one of those terms might be "bird eating insect." Now what happens is Inquire takes that search, and it directly compares it to all 300 million images within seconds. It's engineered to do this, both quickly and efficiently, which is important because it means the system is truly interactive. But it also requires far less computational power than a generative AI approach like ChatGPT. Now once all of these images are sorted based on their relevance to the query, it's really easy for a scientist to just focus their attention on the data that’s most likely to be relevant to them and quickly verify the true matches. Now you have human-verified examples of data that you can directly export and analyze.展望未来:AI助力生态保护
Inquire系统的成功表明,我们可以快速获取隐藏的知识。科学家们已经利用该系统探索了森林在火灾后如何再生,城市和农村地区物种死亡率的差异,以及开花事件与气候变化的关系。这种开放式的系统允许科学家们提出他们感兴趣的问题。未来,我们可以将这种方法应用于生物声学记录、航拍视频、卫星数据和动物项圈的GPS轨迹等各种生态数据类型,从而发现隐藏在不同数据类型之间的联系。虽然AI本身无法解决全球自然危机,但它可以最大限度地利用我们已经收集的数据,帮助我们了解知识差距,并有针对性地收集新数据,从而降低信息获取的成本和时间,并支持保护行动。
Original English
One of our collaborators used this system, and they found thousands of examples of birds eating insects, but also seeds, fruit, nuts, carrion, nectar, plants. And then they took that data that they discovered quickly and they analyzed differences in species' diets between summer and winter. Now what they found was that, yeah, some birds do eat insects in the winter. American robins actually do, but far less than they do in the summer. And some species, like American tree sparrow, that are incredibly dependent on insects as a food source in the summer, don't eat them at all in the winter. This entire process, question-to-answer, took them about three hours. Another team spent 1,560 hours manually curating the data to do a similar study. And when you compare the results from Inquire to that study, you see an almost perfect match. I think this is so exciting, right? It means that we can start quickly getting access to all of this hidden knowledge. And really, I've been so inspired by the creativity of the scientists using the system. All of the flexible ways that people have explored many, many different questions. Things like looking at how forests regenerate after fire. Or discovering differences in species’ mortality between urban and rural areas. Or looking at how flowering events are changing in relation to a changing climate. The possibilities are truly endless. And the fact that it's open-ended means that any scientist can ask the questions they're interested in.共同构建地球生命的完整图谱
我们正处于一个独特的历史时刻,既面临着前所未有的生物多样性危机,也拥有前所未有的应对工具。全球数百万人渴望为自然保护和科学发现做出贡献,而AI工具能够帮助科学家在人类无法单独完成的规模上发现数据中的模式。未来的保护工作不仅在于偏远的雨林或深海海沟,更在于隐藏在我们的生态数据库中,包括我们现在拥有的和尚未收集的数据库。每个人都可以为此做出贡献,上传照片、录制声音、分享观察结果,每一个数据点都是拼图的一部分。我们知道现在必须采取行动来拯救受威胁的自然,并借助科学的AI工具,共同构建地球生命的完整图谱。