With one click, users can hear authentic Cantonese pronunciation through the Cantonese Pronunciation and Lexical Database; refine translations with the Lin Yutang Chinese-English Dictionary of Modern Usage; and explore character origins through the Multi-function Chinese Character Database—where even oracle bone script becomes accessible.

These widely adopted digital language resources were all developed by the Research Centre for Humanities Computing at CUHK. Among them, the Multi-function Chinese Character Database stands out, with over 280 million searches, supporting education, academia, and professional applications. In 2021, Professor Kwan Tze-wan’s integration of this tool with his scholarly work in philosophy of language was awarded the highest rating of 4* (world leading) in the Research Assessment Exercise (RAE), and was later selected as a featured knowledge transfer project— demonstrating the translation of scholarly research into societal impact to meet contemporary needs.

In this issue, CUbicZine interviews Professor Kwan, founder of the Centre and Emeritus Professor of Philosophy at CUHK. After more than three decades of teaching and research in German philosophy, phenomenology, and philosophy of language, he has also pursued a complementary trajectory: translating theoretical inquiry into practical tools, and contributing foundational infrastructure for Chinese language development in the digital and AI era.

He describes the journey as “winding and rugged, but rewarding at every stage.” Through this interview, one begins to sense that behind a clear and intuitive interface lies not only profound philosophical craftsmanship, but also a responsibility to nurture future generations.

The Research Centre for Humanities Computing has developed a number of major databases, including the Lin Yutang Chinese-English Dictionary of Modern Usage (online, 1999), the Lexicon of Confucianism (1997), and Chinese Character Frequency Statistics for Hong Kong, Mainland China and Taiwan – A Transregional, Diachronic Survey (2001), as well as Chinese Character Database: With Word-formations Phonologically Disambiguated according to the Cantonese Dialect (2003). Its flagship project, the Multi-function Chinese Character Database, serves a wide user base beyond academia, with over half of its users coming from non-educational sectors. While early users were mainly based in Hong Kong, recent years have seen significant growth from Chinese Mainland. The latest user distribution: Mainland 51%, Hong Kong 25%, Taiwan 14%, with the remainder from the United States, Southeast Asia and Europe.

From Philosophy to Digital Lexicography

Founded in 1993, the Centre was a pioneer in “humanities computing,” well before the internet became widespread. Kwan began by working with Professor Lau Chong-fuk—then a graduate student—to digitise philosophical classics. Over the past two decades, their focus shifted to Chinese character databases.

Why is it led by a philosophy department? Kwan admits many find it surprising. While rooted in Chinese studies, the work requires a truly interdisciplinary approach, combining linguistics, philosophy, and information engineering. Drawing on Western philosophy and language theory, he re-examined the structure, genesis, and transformation of Chinese script. As the saying goes, support follows a worthy cause. The database project received generous support from colleagues in the Department of Chinese Language and Literature, the Department of Computer Science and Engineering, and the Information Technology Services Centre. And there was also a more down-to-earth reason: “At the time, it just so happened that there was a fool who was willing to devote his time and energy to this project outside of his main job.”

Over the past three decades, the Centre has operated with no recurrent funding support beyond initial seed money. Its main support has come from two Research Grants Council awards and the Hong Kong Quality Education Fund. Despite limited resources—usually fewer than three salaried staff—Prof. Kwan has remained closely involved throughout without additional remuneration. The photo shows Prof. Kwan with key research assistants at different stages of the project. (Source: interviewee)

A Long Commitment to Language Education

What he calls “a fool” is, in truth, an enduring devotion to the nurturing of language.

Since 1999, the Centre has launched a series of online language resources. Among them, the Cantonese database, launched in 2003, has remained widely adopted. Professor Kwan attributes its success to the concept of “phonological disambiguation through provision of word formations”— contextualized pronunciation disambiguation of heteronyms by pairing them with words according to actual usage. Its practical, distinctive design earned the platform the Hong Kong Government’s “Meritorious Website Contest” in 2013.

This Cantonese pronunciation database focuses on heteronyms—particularly context-dependent pronunciations—by pairing characters with contextual word formations to disambiguate their pronunciations. Prof. Kwan noted that Google approached the team twice during development to seek permission to use its lexical pairings data, but no agreement was reached.

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

Around 2005, amid debates over the language of instruction in higher education. Professor Kwan concerned about the long-term development of Chinese, and the preservation of its academic and cultural vitality.

In response, the Centre launched the Multi-function Chinese Character Database project in 2007. By incorporating elements of ancient scripts and focusing on the deeper systematisation of character genealogy and etymology, the project sought to lay essential groundwork for the teaching and study of the Chinese language and its writing system. This time, he devoted more than a decade of sustained effort to the work.

“If something is worth doing, and in the end no one does it, I simply couldn’t live with myself.”

 

The “Kung Fu” Behind Chinese Characters

A decade of effort, honed to mastery. At its core, the logic behind character formation reveals a deeper philosophy of thought.

Take the character (swallow). If it refers to a bird, why is it classified under the “fire” radical? The scholar explains that traditional radical systems—largely inherited from the Kangxi Dictionary—no longer fully reflect how characters originally made sense. Over time, as scripts evolved for easier writing, some of the earlier layers of meaning were lost.

Prof. Kwan explains Lishu-transformation (隸變) as the structural changes that occurred when clerks of the Qin–Han period simplified and modified seal script for practical writing, disrupting original semantic relations embedded in archaic script forms. Take the character 燕 (“swallow”) as an example: in oracle bone script it depicts a bird with a beak, spread wings, and a forked tail; in small seal script, the tail gradually evolved into a form resembling the “fire” radical.

In building this system, the team set out to address these gaps. They drew extensively on pre-Qin sources—especially oracle bone and bronze inscriptions—and looked beyond traditional Sinology, borrowing ideas from Western philosophy to rethink how meaning is constituted. As Kwan puts it, Chinese characters can be understood as expressions of how humans engage with the world, and should be presented in a way that makes that logic visible. His approach draws on Wilhelm von Humboldt’s ideas on language, Husserl’s “pure logic of meaning,” and Merleau-Ponty’s phenomenology of the body.

The Multi-function Chinese Character Database integrates character genealogy, componential trees, and systematic form–meaning interpretation. Its interface systematically presents oracle bone inscriptions, bronze script, bamboo and silk texts, small seal script, Shuowen, and original semantic meanings. Each oracle bone or bronze inscription image can be clicked to trace the provenance of the textual instance (token).

Technically, the project combines this philosophical framework with computing. Characters are represented through componential trees, which map how their parts relate and evolve. This was one of the biggest challenges: interpreting ancient scripts, analysing characters into their components, and reconstructing how they fit together. Rather than being drawn one by one, these structures are generated automatically. Once the main components are defined, the system tracks their relationships in real time and displays them as componential tree structures.

The Multi-function Chinese Character Database was launched in 2014. Even after his retirement in 2016, Professor Kwan remained closely involved in its development, while Professor Lau Chong-fuk took over as director of the Centre. By 2018, the system had reached a mature stage, and it continues to be refined and expanded today—serving as a sustainable knowledge infrastructure for the digital future of Chinese language education.

Crossing Disciplines, Learning Each Other’s Language

Moving from academic world to real-world application, knowledge transfer is not just about integrating technology—it depends just as much on cross-disciplinary collaboration, and knowing when to step back.

His office is filled with books on philology and ancient texts. With a smile, Professor Kwan remarks, “My Chinese is much better now than when I first joined CUHK.” To bring the database to life, he ventured into the worlds of palaeography and coding—areas where he still humbly sees himself as an outsider—yet carried on, tireless and unrelenting. Alongside teaching himself ancient scripts, he traveled to Europe and the United States to attend courses and workshops in information engineering. “Starting to code at my age, I won’t claim to be very good—but at least I’ve learned the fundamentals.” It is precisely this cross-disciplinary grounding that made meaningful collaboration with experts possible.

He also stresses that in the early stages, one must be hands-on. Whenever a new module was developed, he would work through it himself, learning each step in detail—including writing hundreds of form-and-meaning entries. Only by doing so could he fully grasp the practical challenges and guide the team with clarity. Once the project found its rhythm, “you can step back to focus on strategic direction.” On the backend, every entry and revision is systematically recorded, capturing both author and timestamp to ensure transparency and accountability.

To promote the database, the Centre has regularly held talks for secondary school teachers, reaching nearly 200 users. Many teachers find it highly practical: instead of abstract explanations, they can now draw on componential trees, visual forms, and short notes to make character origins much clearer for students. (Source: interviewee)

 

In this way, knowledge is translated from theoretical inquiry into practical implementation—carrying philosophy into the world, advancing the development of digital humanities, and opening new ways for people to understand and live with the Chinese language.

Over three decades, the labour of building and transforming has, in turn, nourished his own thinking. In his eyes, he is its first beneficiary his journey in the philosophy of language still unfolding, still opening onto new understanding.

Last year, Prof. Kwan published Leeway toward Understanding Chinese Language and Script: Philosophical and Cross-Cultural Perspectives, bringing together his highly rated RAE research and reflections from building the Multi-function Chinese Character Database.

Prof. Kwan explains that “Leeway” carries a dual meaning: it signals a new direction in Chinese language and character studies beyond traditional philology, while also reflecting how the database project deepened his own philosophical research—particularly on issues rarely addressed in mainstream philosophy. In recent years, he has lectured internationally, sharing his work on philosophy of language. Photo shows Prof. Kwan speaking in Germany and Taiwan’s universities. (Source: interviewee)

Read More