极速克隆:构建 AI 数字分身
今天我要做一个非常有趣的尝试:在 15 分钟内创建一个自己的视频虚拟化身,并生成一段长达一分钟的宣传视频,主角就是我——Claire Ho。
现在的 AI 产品开发确实不容易,特别是当你需要连接团队和客户依赖的各种工具、设置代理权限、并确保生产环境的可靠性和成本效益时,大部分团队都选择自己摸索,这会耗费大量精力。Merge(用于生产级 AI 的基础设施层)试图解决这个问题,它连接了数千种工具,为代理提供了安全的执行环境,并优化了模型路由和成本。
回到我的实验,我要再次使用我几周前简要介绍过的 Google Flow 和新的 Gemini Omni 视频生成模型。我将竭尽全力创建一个可以进行动画或影视创作的 AI 虚拟化身。根据官方介绍,该功能允许用户创建自己的数字分身。我们之前在它刚上线时尝试过,当时没成功,但这次我们决定再给它一次机会。
创建过程非常迅速:只需扫描 QR 码,通过手机摄像头拍摄几张照片,并按照要求转动头部。整个过程非常流畅,很快系统就完成了扫描。
Original English
Today, I am doing a very strange episode where I'm going to create a video avatar of myself and in about 15 minutes get to a full minute long video starring none other than your favorite podcast host, Claire Ho. Let's get to it. This episode is brought to you by Merge. Building an AI product is one thing. The hard part is everything around it. Connecting to the tools your team and customers rely on, letting agents take action with the right permissions, and keeping everything reliable and cost-efficient once you're in production. Most teams end up piecing that together themselves. So, instead of building the product you actually care about, you get pulled into integrations, permissions, routing, and all the infrastructure underneath. Merge is the infrastructure layer for production AI. It connects to thousands of tools, gives agents secure ways to act inside them, and optimizes model routing and spend without you building or owning any of it. Open AI, Dropbox, and Ramp already use Merge to move fast and build AI right. Visit merge.dev/howiai to start building for free. This episode of How I AI is going to be an adventure because I'm going to be honest, I'm not 100% sure this is going to work. I'm going to return to a product I covered very briefly a couple weeks ago called Google Flow and the new Gemini Omni video generation model and I'm going to try really hard to create an AI avatar of myself that we can animate or I guess cinematically create using AI. So, this is Google Flow and one of the features of Google Flow and the Omni model is you are supposed to be able to create an avatar of yourself. Now, we tried this the day it came out, it did not work, but we're going to give it another college try and see if we can get a full-featured avatar of myself that then we can go and build consistent character videos off of. So, I'm going to select up here. I'm going to create an avatar. We're going to click get started. I'm going to scan this QR code. I have my phone here. I've done this before, so hopefully it'll be fast. Okay, I'm going to put the mic away just for 1 second. I'm going to allow access to my camera, and we're just going to take some photos. Okay, ready? Start. 17 81 49 20 25 22 Okay, now it's having me turn my head. So, I turn my head that way. Give me a check mark. Turn my head the other way. It is giving me a check mark. And and it says we're done. Now, it said we were done last time we tried this. So, we're going to see. It's going to take a couple minutes, and then we will come back and see if I can actually use this avatar of myself.
创意赋能:AI 驱动的创作流
几分钟后,我的虚拟化身生成了。虽然是一个略带“鱼眼镜头”感的版本,但我现在可以利用它来为“How I AI”播客创建一个宣传视频。我进入 Flow,输入我的需求:“帮我为 How I AI 播客创建一个宣传视频故事板。我已经有一个名为‘me’的化身可以引用,请帮助我构思几个精彩的场景。”
我最喜欢 Flow 的一点是,它不仅是一个视频生成工具,更是一个完整的创作套件。它不仅能帮助我生成化身的动画,还能辅助我进行视频的头脑风暴。我很有创意,但并不精通视频制作,所以我非常期待它能帮我搞定这一切。
我构思了一个在黑暗家庭办公室、背景有 AI 书籍和海报的场景,具有一种黑客风格的真实生活感。尽管我有不少输入错误,但系统仍然理解了我的意图。生成式 AI 模型(尤其是这些多模态模型)最吸引我的地方在于,它让我能够独立制作以前绝对做不到的内容。我原本很难构思框架、进行分镜,但现在我有了一个 AI 制作人来协助我完成这些任务。
经过几次尝试(中间因为选错了生成模式导致生成了图像而不是视频,稍微调整了一下),它很快生成了一系列引人入胜的视频片段。比如我操作机械键盘的特写、办公室全景、我在人体工学椅上的旋转以及数码平视显示器的特效。我很快就把这些片段拼接在浏览器中的编辑时间线上。
Original English
Okay, so look at this beauty. Um there's this fish eye lens version of me that is now an avatar. So, I supposedly can use this. And let's use it to create a hype video for the How I AI podcast. So, I'm going to go in here and say, "Help me create a storyboard for a hype video for the How I AI podcast. I already have a character named me we can reference. Help me come up with the few scenes that would make this great. This is a podcast by Claire about the best ways to use AI at work and in life. Exclamation mark. Okay, so what I love about Flow, or what I was pitched to me about Flow, is that it's not just a video generation tool, it's actually a whole creative suite. And so, ideally it's going to be able to help me not only animate or video generate this avatar of myself, it's also going to help me actually brainstorm what this overall video could be. And I'm you know, I'm creative, but I'm not video creative, so I'm excited to see what it looks like. So, how do you imagine Claire? Is she in a modern studio or perhaps a bright airy home office? Should it feel high-tech and sleek or more grounded and lifestyle focused? And are we going for high energy and fast-paced and thoughtful and inspiring? So, I'm going to say she is in a dark home office, dark green walls with books about AI and fun posters, lighting around. This should be more authentic lifestyle version but it's high-tech and about coding. Have a hacker vibe to it. Okay. Well, I have a bunch of typos, but we'll see we'll see what this does. And what I love about these video models and these new tools, again, usually here on How I AI we talk about coding, we talk about website generation, we talk about PRDs and work product. But what I really appreciate about these new generative AI models, in particular these multimodal ones, image and video, is it unlocks for me an ability to generate create something that I would have never been able to do before. So, I would have never been able to um solo produce a high video for my podcast. I would have a hard time brainstorming it, I wouldn't know how to frame it, I wouldn't know how to block it. But now I have this AI producer here that can help me with this effort. So, let's see what the frames are. It's about seven frames. Um it's going to be an extreme close-up of me typing on a mechanical keyboard, totally on brand. Um, then there's going to be a wide shot of the office, then it's going to reveal me in my ergonomic chair. Spoiler alert, I'm not actually in an ergonomic chair. I'm going to spin around, that's going to be funny, and it's going to give me a digital heads-up display, which is also ridiculous, but let's let it happen. Then it's going to do a very, what I'm presuming to be a very cheesy AI montage, a lifestyle moment, a call to action, going to hit you with the the podcast uh microphone, and then it's going to say how I AI. Um, if this looks good, I'm going to say, "This is great. Generate the storyboard. I already have the character at me." Um, and so I'm going to send that. We're going to see what it comes up with. I've noticed that it has a hard time referencing the me character in some early tests, so let's see what it comes up with. I'm presuming it's going to take a couple minutes, so we will take a mini break and then come back to see what it looks like. Okay, it looks like it's generating a uh grid for the storyboard. It can't use the avatar, so I think it's going to do it without the character reference. It'll be really interesting to see what it comes up with. But then as soon as it's ready, I'm going to go ahead and generate at least a couple of these storyboard scenes one by one, and we can see how well it does with my avatar. Oh, I mean, this is delightful. Look at this glowy mechanical keyboard. Look at how I am hacking on three keyboards. I'm going to make a little eyes at you with my my fake glasses, my very trendy glasses. There's going to be me dragging and dropping a file that probably says like ai.md. I'm going to smile, and then I'm going to speak into the podcast. "This looks great. So, what I think I'm going to do is I'm going to paste in this first frame of the video that the agent came up with, and instead of saying Claire, I'm just going to at mention in this avatar that it gave me, um so that we can see if it generates this video with me as the character. And so, I think I've replaced my name here. Um I've given details on camera, on lighting, on everything. I press enter. Let's see what it creates with my avatar. I have no idea what we're going to get into, and hopefully it won't be terrifying. Okay, I'm already nervous. What is surprising to me that I didn't actually expect is it does have my posters and my books background here, I guess because they're behind me when I took the photo, it's taking advantage of that. And I'm going to share my audio as well, and we're going to see how this video worked. Okay, I got that wrong. I actually generated images instead of videos. Totally messed up, did not click the right thing down here in the bottom right. I had the image generation instead of video generation. So, again, I'm going to paste that walk-through of the scene here. I'm going to replace my name with the me avatar. It's going to have my fingers flying across that mechanical keyboard. It's going to be so cool. I'm going to go ahead and press send, and we're going to see how long it takes to generate a video. Now, something you'll notice about every time you generate videos and it used to work like this in Veo 2, so I'm not Veo 3 as well. So, I'm not surprised they do this as they're generating two versions of it. It's going to take a couple of minutes. The image took a couple seconds. These are probably going to take a couple minutes, so I will come back, and hopefully we will have our first video with Claire's face in it. And while we're waiting, I'm going to queue up one or two other scenes, and see if we can get ones going with my actual face in it because some of these had um like the back of my head as opposed to my face and I think we want to see what my face avatar looks like. So, we'll pick frame three and see if we can get that going as well.
成果与局限:跨越恐怖谷
这次实验不仅迅速,而且结果超出了我的预期。虽然并非 100% 完美,但我认为它已经达到了 90% 的完成度。
从个人特征来看,它大约有 50% 的时间看起来很像我,另外 50% 则处于一种微妙的“恐怖谷”状态,尤其是当涉及情绪表达(比如大笑场景)时,看起来会有些奇怪。此外,存在一些角色一致性问题:背景、头发样式和光线会随场景变动。它有时会抓取到拍照时背景中的元素(如海报、书籍),这些元素在不同片段中会发生变化。
在内容表现上,视频生成器对 AI 和高科技的理解依然带有 2000 年代的刻板印象——比如我在视频中拿着一个 24 英寸的 iPad,看着一个看起来像教堂的示意图,这非常令人困惑。
尽管如此,考虑到我从零开始只用了大约 15 分钟就生成了这个一分钟的视频,我感到非常震惊。这不仅仅是一个爱好,我认为如果进行一些持续的背景提示调整,投入一点点更多的时间和额外的素材,我完全可以制作出一个足以瞒过大多数人的宣传视频。
Original English
Okay, the first video generated. Now, we have blue nail polish. I still like it. Okay, let's see. We were told AI would replace us. >> [laughter] >> That is quite spooky. Okay, we were told AI is going to replace us. Let's see if the video with me actually generates a callback to that. Um so, while that's generating, I'm going to go ahead and make all of these. We're going to stitch them together. It's going to be so awesome. So, stick with us. We're going to generate a bunch of videos and we're going to stitch it together into one long hype video. This episode is brought to you by Jira Product Discovery. AI has made individual PMs incredibly productive, but multiplayer mode is where it still breaks. Getting everyone aligned on what should actually get built. Decisions live in a markdown file from last week. The road map's a spreadsheet no one's looking at. Jira Product [music] Discovery is where teams actually decide what to build. Capture ideas, prioritize them as a team, and share a living road map everyone works from. It's powered by Atlassian's team work graph, so it can pull in customer feedback, what your team shipped, plus your goals, and suggest what to build next. And when a decision is made, you can hand it off straight to Jira, so a developer or even an agent can pick it up and start building. Teams at Canva, Deliveroo, and Toast already use Jira Product Discovery. Join more than 25,000 teams at at lassian.com/howiaai. Start building the right things together. Okay, I have 17 generating but while we're waiting for those to finish I just cannot oh >> [laughter] >> Sorry. Sorry for you all that are listening and not watching. I just got um jump scared by the AI version of myself wearing glasses um turning around in a spinning chair. So take a look at both of these. This one's pretty good. I'm spinning in a circle. Okay, sorry. Back to those I need to describe this for. Um So this is using an AI avatar of myself. The prompt was I spin my ergonomic chair around to face the camera. I push my glasses which I don't have up to the bridge of my nose and I say this is Claire. I am Claire and this is how I AI. Let's watch V1 of this video which is actually a scream riot. I'm Claire and this is how I AI. >> [laughter] >> Okay, it was actually pretty good. Um what's really funny is I do have the it has the Nvidia way in the background which I don't have right here but I do have upstairs so I do believe the AI overlords are really paying attention. Um I want to make you laugh and look at the second version where I spin in a circle twice pretty good. I'm Claire [music] and this is how I AI. This one got my um not curled hair a lot better but I prefer the other video. It makes me look a little bit nicer. Okay, I'm going to take 1 minute. I'm going to stitch all these videos together in the form factor that Gemini told me I should that flow told me I should. We're going to bring this how I video together. I'm going to show it to you end to end and then I'm going to conclude today's very strange episode of How I AI where I use my avatar to create an end-to-end hype video for this podcast. Cool. So, it actually seems like I can show you a little bit of how we're going to stitch this video together. So, if you see here, once I click into any one video, I have a video editor timeline here that I can use right in the browser to stitch together all these videos. So, I'm going to go ahead and add these in the order that the original AI told me my hype video should go and then we'll look at it end to end and we'll see if we really like it. Okay, this took me about 5 minutes, but all I did was um stitch together my favorite versions of all these avatar-generated AI videos um scene by scene about seven of them together um to show one end-to-end hype video. Again, this episode is probably going to be sub 15 minutes. That includes recording my face as an avatar, figuring out what the heck is going on with this tool, building a storyboard, generating all the videos, and stitching them together here in this editor. And now, the worldwide debut of the How I AI hype video. I am going to show you who knows what we're about to get, but we're about to get it. Here we go. We were told AI would replace us. My god. [laughter] >> I'm Claire and this is How I AI. From automating the mundane to dreaming up the impossible. >> [laughter] >> And it's about the tools that change the way we live and work. >> [laughter] >> Join me as we deconstruct the future one prompt at a time. Subscribe to How I AI. >> How I AI. Available now everywhere you Available now everywhere you get your podcasts. >> Okay. I am actually obsessed with this. Let's talk about what I love and what I don't. What I love. This took zero time and effort. And it is I wouldn't say it's like 80% there, but is it 50% there? 100% yes. Am I going to tweet this? Immediately. Absolutely. Did this take no effort? Basically no effort, no knowledge. Okay. So, what did I like about this avatar experience? You know what? This is like kind of my face. It's not quite my face. I would say about 50% of the time it's my face and 50% of the time it's like uncanny version of of my face. Some things I noticed from a character consistency perspective. This gave me beautiful long wavy hair which I have recently cut off cuz I have a child. So, you see there's like a location inconsistency. This background has has books and a um an hourglass. This background is a different color and it has plants. It pulls in some things from my avatar like it pulls in this poster that was in the background of when I took my photos and it changes a little bit over time. And so, you can see the books on the shelf change, the lighting changes. As always, these video gen and image gen models are really early 2000s coded on what they think AI and impressive technology is. So, I'm holding like a 24-in iPad in this video looking at a schematic of It looks like a church. It's very confusing. The heads-up display that shows up on my face when I'm looking at AI. I'm I'm apparently um coding in in Gemini a robot of some sort. So, it's pretty hilarious, but even looking at this frame, I would say this is the one that felt like it looked most like my face. Like I'll just try to look serious so you all can see. It's pretty good. It's got It's even got my sun damage here. So, good job, Gemini, not um smoothing out smoothing out my face. And so, I do think this is 90% there. Um not 100% there, but it's really interesting even um seeing my face turn left and right how accurate it got on the side profiles of my faces. Now, this scene right here where I'm laughing 100% uncanny valley. I look very strange like I'm on some side of um medication perhaps. And so, I'm not sure it 100% has emotions really well and some of the timing and hiccups you notice while you were watching the video. I I spoke over myself. Those sorts of things, but this scene right here is legitimately pretty good. I bet with some um consistent background prompting, with a little bit more effort, with some additional images going into this omnimodel, I think I can make a hype video that would convince most of you if not all of you. Now, do I think it's great at typography? Do I think it's great at graphics? No, this is kind of lame. This ending part is kind of lame. But again, we're talking probably 10 minutes top to bottom. So, we're we're talking, you know, probably 15 minutes from very beginning, knew nothing about this tool, to I have this 1-minute video now I can share with you all. I'm pretty blown away, you guys. And so, I'm going to go spend a little bit more time with the Google omnimodel. I'm going to spend a little bit more time with Flo. This might be my new favorite hobby project. I'm kind of obsessed with it. I want to hear if you all are willing to put your avatar in here, if you can actually get it to generate consistent characters, and what your experience is using these kind of incredible new video models. So, I know this is a little bit of a different style of how I AI. We usually do coding, we usually do work stuff. This is a tool I did not know. This is a process I'm very unfamiliar with, and I really think I got an outcome that was much better than I expected with very little knowledge of the tool. So, if that is not a How I AI success story, I'm not sure what is. I hope you enjoyed this very strange mini episode of How I AI. I cannot wait to see what you generate, and please share your examples in the comments. Thanks for joining. Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at how I AI pod.com. See you next time.
📌 文中提及的人物和组织
公司/组织: Google, Merge, Atlassian
产品/模型: Gemini Omni, Google Flow, Jira Product Discovery