I lead innovation and build the platforms, teams and governance that take AI from idea to audited, scaled deployment across regulated healthcare and pharma R&D.
I’m an AI and computational-sciences leader who turns frontier models into validated production systems for regulated healthcare, pharma R&D and drug discovery, across research, product, engineering and governance.
Built and scaled AI teams (4 → 15+, ~25 across the org) and the platforms behind them, with NVIDIA, Microsoft, OpenAI and AWS.
Cut validation cycles from 40 weeks to one, accelerated biologics screening, and shipped audit-grade GenAI in regulated settings.
Operate across scientists, engineers, vendors and executives, with Responsible-AI governance to EU AI Act standard.
Figures are representative of internal programme outcomes; details simplified for confidentiality.
2 US (behaviour-signal & data-perspective, Optum), 11 Singapore (B2B prospecting, Leadbook) and 2 at Nokia Bell Labs; peer-reviewed at AAAI, DASFAA and IEEE Big Data, spanning topic modelling, multimodal indexing and privacy-preserving clinical NLP.
Novartis Galaxy Team Award three years running (Biologics AI 2025, AI4Biologics 2024, Horizon / PKS AI platforms 2023), plus the Star Award for Leadership (2022): four consecutive years recognised for technical leadership. Earlier: AAAI-2014 Scholarship, the Extra Chapter Challenge Award (NUS Enterprise · PhD commercial-feasibility fellowship), the Google Developer Challenge Scholarship (Udacity · Google) and Winner, Startup Weekend Singapore.
An integrated antibody-discovery platform built around a reinforcement loop. Large-scale molecular-dynamics simulations are orchestrated across ~80 GPUs on DGX; their trajectories feed feature extraction that fine-tunes ML emulators of biophysical and developability properties — stability, aggregation, manufacturability. Predictions are benchmarked against wet-lab results (lab-in-the-loop), and the gap drives the next round. The emulators cut compute cost while preserving accuracy, alongside de novo design and binding-affinity optimisation.
An end-to-end multimodal platform that automates medical-claim and material validation. The hard part was the inputs: messy, unstructured lab reports with embedded charts and tables. Custom-built models parse those charts and tables into structured fields, a library of ~20 specialist models is orchestrated over them, and a unified evaluation layer with human-in-the-loop gates and full audit trails turns a multi-week manual review into a near-real-time, traceable workflow.
PK assay reports hide their numbers in some of the hardest tables in pharma: multi-page, side-by-side, paragraph-shaped, with grouped keys (Fu mean / SD) and units that sit in a different cell. I reframed extraction as a spatial search problem. A custom layout model finds the tables and columns, then a graph method walks the natural reading order to pair every key with its value. It is config-driven and unsupervised, so scientists define new extractions themselves: no labelled data, a confidence score on every page, and under a second per page on a laptop. It outperformed heavyweight commercial and deep-learning extractors that needed thousands of labels and still missed complex layouts.
A first-of-its-kind agentic system for continuous, autonomous R&D, running as an iterative loop: ideation, then agents that build their own library of papers and tools end-to-end, exposed through a library agent over tools and MCP that both other agents and human researchers can query. Specialist personas — orchestration, planning, search, code, critique, writing — run across 10+ heterogeneous LLMs, with shared memory, human checkpoints, GPU/cloud orchestration, in-silico benchmarking and a comprehensive report after every cycle.
AI-driven therapeutic-target identification that integrates multi-omics and high-dimensional biology with graph-based learning to surface and rank novel targets. A companion biomarker-discovery effort (imaging + proteomics) advances predictive signatures across neuroimmunology programmes, including progressive multiple sclerosis.
A multi-agent RAG system that mines and validates immunogenicity evidence from VH/VL sequence-level signals all the way to clinical findings, FDA approvals and labels, patents and published studies. Extractor agents pull candidate evidence; critic agents challenge and cross-check it; a validation step enforces structured, dataset-level QA with traceable provenance — feeding downstream developability and risk assessment with evidence you can audit back to source.
A sibling to Horizon MAP that turns the same governed foundation toward creation: an LLM system that generates marketing and communication materials — copy, layout and on-brand visuals — from a short brief and a set of brand and compliance rules, with review built in. The same orchestration and guardrails that validate documents are reused to produce them.
A system that transforms raw claim data into temporal signals using a custom-built embedding model, then tracks how those embeddings move through space and time across different layers of the healthcare system — facility, provider, member. Where an embedding drifts away from its neighbours, a previously invisible medical-spend driver surfaces — turning claims into prioritised, explainable affordability levers, with dynamic dashboards to explore them.
An engine that automates medical cost-saving ideation. It reads clinical and claims signal across the population and translates it into a ranked set of actionable, explainable affordability levers — surfacing where care can be delivered better and cheaper, and handing analysts a prioritised worklist instead of a blank page.
A high-throughput assistant for claim review and payment integrity. NLP reads each claim, an explainable-AI layer surfaces the reasons behind every call, and the system recommends pay-or-review under strict audit and compliance constraints — keeping a human in control while clearing the routine volume fast.
A large-scale smart news-aggregation system built in the lab: it ingests news as text, image and video, classifies it with multimodal models, and organises everything under shared, machine-generated topic labels so the whole corpus becomes searchable and linkable across formats. The indexing core (ANNOTATE) was demonstrated at Mobile World Congress 2018 and published at IEEE Big Data.
Temporal models for passive human–computer interaction from EEG, ECG, EMG and eye-tracking signals. The flagship demo infers intent from gaze: you think of an object and stand before a gaze-tracking screen; as images cycle, the model reads where your eyes settle and walks down the ImageNet hierarchy — narrowing from broad categories to specifics — until confidence passes 80% and it returns its top-5 guesses. No clicks, no typing; intent inferred from gaze.
A scalable pipeline that turns raw video into structured, multimodal object chunks. It ingests media at scale, segments each video, and runs fast deep-learning models to label every segment — finding the inner similarities and relations across a library so that video becomes searchable, linkable content rather than an opaque stream.
A system to represent, infer and communicate meaning for personalised content. Unsupervised hierarchical topic models — with experiments in deep generative networks — learn a compact representation of meaning that can be transferred and re-expressed per user, with evolving disambiguation of senses and generalised topic labelling across multimedia, so the same message adapts to each recipient.
One of Asia’s largest B2B intelligence graphs — tens of millions of verified company and contact records merged from across the web — with a prospect recommender built on a patented Company–Product–Customer “deep relationship” model that learns which new prospects resemble a customer’s best existing ones. The engineering ran from distributed crawling and entity-matching to real-time lookup.
Doctoral work that takes millions of flat user comments and posts and turns them into a temporal aspect–action graph: a joint aspect–action topic model infers what people are talking about and what they intend to do — without labelled data — and arranges it into a structured, time-aware hierarchy of who said what about which aspect, when. Published at AAAI; the foundation of a self-supervised discussion-analysis and prediction system, supervised by Prof. Chua Tat-Seng.
A comprehensive AI platform for legal teams — six connected modules that carry a matter from intake to defensible output, with cited reasoning and an audit trail running through all of it.
A physics-based, pixel-level OLED guardian that protects against burn-in — and keeps working inside games by automatically detecting static HUD elements and treating them in real time.
A VR chemistry-lab training simulation, prototyped with a custom liquid-interaction engine and a library of lab protocols — so learners can pour, mix and run procedures with believable fluids in a safe virtual lab.
A personal operating system of AI agents that helps you run your life — capturing, planning and offloading the mental tasks you'd otherwise juggle, so more of the day's overhead runs itself.
A system that helps you get 1% better across the parts of life you care about — small, AI-guided adjustments, tracked over time, compounding into real change.
A “Wikipedia of processes” — one place gathering the steps for anything with a procedure: applying for a job, onboarding at a company, a visa application, even becoming an Olympic medallist — structured so you can follow, fork and improve them.
A wearable smart ring that fuses on-device AI vision with ultrasonic / sonar ranging to perceive obstacles and open space, guiding blind and low-vision users with real-time directional haptics. Prototyped in Singapore.
A crowd-sourced, location-based mystery-shopping platform that matches tasks to the right people and places using contextual topic models — everyday shoppers as a distributed sensing network. Winner, Startup Weekend Singapore.
Point a phone at a supermarket shelf and a vision-and-health model recognises each product and paints a personalised health “hue” over it. Runs on the first structured database of Singapore food labels, which we built and processed to train it.