AI can write Filipino. Ask it to build sumulat and it says sulumat.
Three separate tests of Filipino found the same gap. The chatbots sound fluent, then break word rules that Filipino kids follow without thinking.
Ask a chatbot to write in Filipino and it will hand you a clean paragraph in seconds. Researchers have started checking whether it actually understands what it wrote, and the answer so far is no. A test called PACUTE gave AI models the root word sulat, then asked them to insert um, the small piece Filipinos drop inside a word to say it was done. The correct answer is sumulat. One model put the piece in the wrong slot and produced sulumat, a word that does not exist.
PACUTE runs 4,600 of these small tasks. It checks whether a model can pull a word apart, repeat the right syllable, and hear the difference stress makes. Filipino leans on all three. Takbo becomes tatakbo when the running has not happened yet. Basa means read or wet depending on where the stress lands, and almost nobody writes the marks that would tell you which. A Filipino speaker sorts this out without pausing. The models are guessing.
Why it matters
This is the part that reaches your phone. Filipinos already lean on these tools for schoolwork, captions, and translating messages, and fluent output is exactly what makes a wrong answer hard to spot. FilBench, a separate Filipino test published last year, ran 27 of the leading AI models through Filipino, Tagalog, and Cebuano tasks. The best score belonged to GPT-4o at 72.23 percent. Its weakest areas were reading comprehension and translation, the two things most people are actually using it for.
The reason is not that Filipino is unusually hard. It is that there is not much of it in the pile of text these systems learn from. Batayan, a Filipino test published in the 2025 proceedings of the Association for Computational Linguistics, landed on the same explanation: complex word building plus thin representation in training data. Three research teams, working separately, found the same hole.
The catch to watch
These are lab tests, not everyday chat, and a model that fumbles sulumat can still write a decent Facebook caption. The scores also move as new versions ship, so any number here is a snapshot. The useful takeaway is smaller and more practical: the tools are genuinely helpful for brainstorming, rough translation, and sorting your notes, but treat Filipino output as a draft, not a finished answer. If it sounds right, that is not yet proof it is right.