Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
Falcon-Emirati:当大语言模型学会方言、文化与细微差别
Arabic is really a family of languages living under one name. Modern Standard Arabic is what you read in the news or a textbook, but it’s rarely how people actually talk to each other. In the UAE, day-to-day conversation, humor, negotiation, and storytelling happen in Emirati Arabic, a Gulf dialect with its own vocabulary, its own rhythm, and a culture wrapped tightly around it. 阿拉伯语实际上是一个共用一个名称的语言家族。现代标准阿拉伯语(MSA)是你从新闻或教科书中读到的语言,但人们在日常生活中很少这样交流。在阿联酋,日常对话、幽默、谈判和讲故事都使用阿联酋阿拉伯语——这是一种拥有独特词汇、节奏以及深厚文化底蕴的海湾方言。
Emirati poetry, especially nabati poetry, along with proverbs and short anecdotes, carries meaning that doesn’t survive a literal, word-for-word reading. A model that only knows MSA can translate every word of an Emirati sentence and still miss what it actually means. That’s the gap Falcon-Emirati-7B is built to close. It’s a dialect-specialized model on top of Falcon-H1-Arabic, aimed at understanding and generating Emirati Arabic the way a native speaker would: the vocabulary, the tone, and the cultural context behind it. 阿联酋诗歌(尤其是纳巴蒂诗歌)、谚语和短篇轶事所承载的意义,无法通过逐字逐句的直译来传达。一个仅掌握现代标准阿拉伯语的模型,即使能翻译出阿联酋语句中的每一个词,也可能无法理解其真实含义。Falcon-Emirati-7B 正是为了填补这一空白而构建的。它是在 Falcon-H1-Arabic 基础上开发的方言专用模型,旨在以母语使用者的方式理解和生成阿联酋阿拉伯语,涵盖其词汇、语调以及背后的文化语境。
Built on Falcon-H1-Arabic
基于 Falcon-H1-Arabic 构建
We didn’t start from scratch. Falcon-Emirati-7B is built on Falcon-H1-Arabic, our Arabic model family that already set new benchmarks for the language earlier this year. Falcon-H1-Arabic uses the Falcon-H1 hybrid architecture: State Space Models (Mamba) and Transformer attention running in parallel inside every block, with their outputs fused before each block’s projection. That combination gives the linear-time efficiency of Mamba on long sequences while keeping the precision of attention for long-range dependencies, which matters for a morphologically rich language like Arabic. 我们并非从零开始。Falcon-Emirati-7B 构建于 Falcon-H1-Arabic 之上,这是我们今年早些时候推出的阿拉伯语模型系列,已为该语言设定了新的基准。Falcon-H1-Arabic 采用了 Falcon-H1 混合架构:在每个模块内部并行运行状态空间模型(Mamba)和 Transformer 注意力机制,并在每个模块投影前融合它们的输出。这种组合既赋予了 Mamba 在处理长序列时的线性时间效率,又保持了注意力机制在处理长距离依赖时的精确度,这对像阿拉伯语这样形态丰富的语言至关重要。
The family spans three scales (3B, 7B, and 34B parameters) with context windows up to 128K and 256K tokens, and it was already trained on a broad mix of MSA and dialectal Arabic (Gulf, Levantine, Egyptian, Maghrebi) alongside English and multilingual data. That gave us a strong starting point: a model that already understood Arabic broadly, handled long context well, and had some dialectal exposure baked in. 该系列涵盖三种规模(30亿、70亿和340亿参数),上下文窗口高达 128K 和 256K token。它此前已在现代标准阿拉伯语和多种方言(海湾、黎凡特、埃及、马格里布)以及英语和多语言数据组成的广泛混合语料上进行了训练。这为我们提供了一个强大的起点:一个已经广泛理解阿拉伯语、能很好地处理长上下文,并具备一定方言基础的模型。
Falcon-Emirati-7B takes that foundation and pushes it specifically toward the Emirati dialect, the vocabulary, the grammar, and the cultural knowledge that a general Arabic model, however capable, doesn’t pick up on its own. We built Falcon-Emirati-7B on the 7B variant specifically. It’s the sweet spot in the family: large enough to hold onto the nuance that dialect adaptation needs, but small enough that both training and inference stay practical. Falcon-Emirati-7B 在此基础上,专门针对阿联酋方言、词汇、语法和文化知识进行了强化,这些内容是通用阿拉伯语模型即便能力再强也无法自行习得的。我们专门基于 7B 版本构建了 Falcon-Emirati-7B。它是该系列中的“黄金比例”:既足够大以容纳方言适配所需的细微差别,又足够小以确保训练和推理的实用性。
Why Dialect Adaptation Is Hard
为什么方言适配如此困难
Turning a general Arabic model into an Emirati-dialect specialist sounds like a smaller job than building the base model in the first place. It isn’t. A few things make it genuinely difficult: Emirati is mostly a spoken dialect. It shows up far less in writing online than MSA, or even other Gulf and Levantine dialects, so there just isn’t as much raw text to learn from. Meaning is often non-literal. Idioms, proverbs, and poetic references lean on shared cultural context, not surface vocabulary. 将通用阿拉伯语模型转化为阿联酋方言专家,听起来似乎比构建基础模型要简单,但事实并非如此。有几个因素使其变得非常困难:阿联酋语主要是一种口语方言。它在网络书面语中的出现频率远低于现代标准阿拉伯语,甚至低于其他海湾和黎凡特方言,因此可供学习的原始文本并不多。其含义往往是非字面的。习语、谚语和诗歌引用依赖于共享的文化语境,而非表层的词汇。
There’s no established playbook. There isn’t a well-documented recipe for how much dialectal data is enough, how to mix it with MSA and general Arabic, or which training stage (continued pre-training, SFT, or preference optimization) matters most for picking up a dialect. That last point shaped how we worked. A lot of building Falcon-Emirati-7B came down to trial and error: testing different data mixes, training stages, and supervision strategies, and using both human judgment and benchmark scores to figure out what actually moved the needle. 目前还没有现成的操作手册。对于需要多少方言数据才足够、如何将其与现代标准阿拉伯语混合,或者在哪个训练阶段(持续预训练、SFT 或偏好优化)对掌握方言最重要,目前尚无明确的方案。最后一点决定了我们的工作方式。构建 Falcon-Emirati-7B 的过程很大程度上依赖于反复试验:测试不同的数据组合、训练阶段和监督策略,并结合人类评估和基准测试分数,来找出真正有效的改进方法。
Our Approach to Data
我们的数据处理方法
We built a dedicated Emirati data pipeline on top of Falcon-H1-Arabic’s pretraining, drawing on three complementary sources. 我们在 Falcon-H1-Arabic 预训练的基础上,构建了一个专门的阿联酋数据流水线,利用了三个互补的数据源。
-
Authentic Emirati-Dialect Web Data: We crawled and curated content from Emirati websites and forums written natively in the dialect, not translated or transliterated from MSA. This is where we got our ground truth: how Emiratis actually write and speak online, the everyday phrasing, the colloquial expressions, and the natural back-and-forth between Emirati and MSA that shows up in real usage.
-
真实的阿联酋方言网络数据:我们从阿联酋网站和论坛中抓取并整理了以该方言原生书写的内容,而非从现代标准阿拉伯语翻译或转写而来的文本。这是我们获取“真理”的来源:阿联酋人在网上真实的写作和交流方式、日常用语、口语表达,以及在实际使用中阿联酋方言与现代标准阿拉伯语之间的自然切换。
-
MSA Data About Emirati Culture and Identity: Alongside the dialectal text, we pulled in MSA-language material specifically about Emirati culture, heritage, and language: articles and references on local customs, values, history, and social norms, including how Emiratis are perceived and stereotyped. This doesn’t teach the model to write in dialect, but it teaches the model what it’s talking about when Emirati topics come up, things like heritage, etiquette, and the context a native speaker just knows.
-
关于阿联酋文化与身份的现代标准阿拉伯语数据:除了方言文本,我们还引入了专门关于阿联酋文化、遗产和语言的现代标准阿拉伯语资料:关于当地习俗、价值观、历史和社会规范的文章和参考资料,包括阿联酋人如何被看待和刻板印象化。这虽然不能教模型用方言写作,但能让模型在涉及阿联酋话题时,理解其背后的遗产、礼仪以及母语使用者所熟知的语境。
-
Synthetic Data, Guided by Glossaries and Style Rules: Authentic dialectal text alone wasn’t enough to cover the range of topics a chat model actually needs to handle day to day. So we generated a large amount of synthetic Emirati-dialect data to fill the gaps. We didn’t just let a generator model improvise in “Gulf-ish” Arabic. We constrained it with strict rules and glossaries and dictionaries built specifically for Emirati vocabulary and grammar. Those guardrails made the difference between synthetic output that reads as authentically Emirati and output that’s grammatically fine but sounds off to anyone who actually speaks the dialect.
-
由词汇表和风格规则引导的合成数据:仅靠真实的方言文本不足以覆盖聊天模型日常需要处理的各种主题。因此,我们生成了大量的合成阿联酋方言数据来填补空白。我们没有让生成模型随意使用“海湾风格”的阿拉伯语,而是通过专门为阿联酋词汇和语法构建的严格规则、词汇表和字典对其进行了约束。这些护栏确保了合成输出既能读起来像地道的阿联酋语,又不会出现语法正确但让母语使用者感到别扭的情况。