NSF III-2312991

Machine-learning tools have become ubiquitous in modern information systems. The data inputs used by these tools often originate from relational databases. The data outputs generated by those tools are often stored in databases, where they can be used for subsequent data analysis. Typically, however, the learning process itself is performed outside the database system. This project investigates the opportunity for performing more of the machine learning work within the database itself, avoiding expensive (and often redundant) data export and import. In partnership with researchers from Relational-AI and Microsoft, the Columbia University team will design and build two interacting open-source systems named MARQUE and ZORK. These systems will make data analysis more efficient and effective for database-resident information. Improved efficiency will lead to faster, more cost-effective machine learning, and executing ML within the DBMS will simplify operational complexity and benefit from DBMS features such as scalability, access control, and data management. Ultimately, this work will broaden the adoption of machine learning technologies in a wide range of data-intensive disciplines.

MARQUE will be a database management system that supports machine learning primitives such as linear algebra operations within the context of a query processing engine. The system will efficiently compile SQL queries using embedded machine learning models within the database, combining state-of-the-art query processing techniques with highly engineered linear algebra algorithms. MARQUE will allow components of the machine-learning pipeline itself to be formulated as in-database operations, avoiding unnecessary data copying. Conventional SQL analytic queries that can be reformulated using extensions of operators like matrix multiplication can be optimized to use efficient execution plans involving specialized algorithms for such operators. To further support in-database machine learning, the project investigators will build ZORK, a system to support machine learning at scale that will make use of the infrastructure provided by MARQUE. ZORK will scale to very large datasets by processing factorized representations of the data rather than explicitly materializing large joins. This project will develop new and innovative query processing techniques for queries involving both conventional relational operators and generalized linear algebra operators. Tight integration will facilitate query optimization within and between operators. Using this system, a range of machine learning techniques will be developed that operate entirely within the database management system, avoiding data export and simplifying concerns such as data privacy administration.

Principal Investigators

Open Source Software

Publications

  1. SPALM: A Sparsity-Pattern-Adaptive Library for Matrices
    Junyoung Kim, Kenneth Ross
    SIGMOD 2026
  2. Data Flow Control: Data Safety Policies for AI Agents
    Charlie Summers, Eugene Wu
    ArXiV 2026
  3. SANA: What Matters for QA Agents over Massive Data Lakes?
    Austin Senna Wijaya, Jiaxiang Liu, Haonan Wang, Eugene Wu
    arXiv 2026
  4. Law-Abiding Agents: Demonstrating Safe Tax Preparation with Data Flow Control
    Charlie Summers, Prajwal Raghunath, Zhibin Shen, Hardik Gupta, Eric Choi, Mayur Kulkarni, Sally Go, Peter Yu, Akriti Agarwal, Zhuo Zhang, Oliver Kennedy, Eugene Wu
    preprint 2026
  5. LAKEQA: An Exploratory QA Benchmark over a Million-Scale Data Lake
    Haonan Wang*, Jerry Liu*, Yurong Liu, Austin Senna Wijaya, Tianle Zhou, Eden Wu, Yijia Chen, Wanting You, Reya Vir, Daniela Pinto Veizaga, Grace Fan, Yusen Zhang, Juliana Freire, Eugene Wu
    ICML 2026
  6. BranchBench: Aligning Database Branching with Agentic Demands
    Elaine Ang, In Keun Kim, Sam Weldon, Kevin Durand, Kostis Kaffes, Eugene Wu
    ArXiv 2026
  7. BranchBench: An Extensible Benchmark for Agentic Database Branching
    Elaine Ang, Kostis Kaffes, Eugene Wu
    CAIS Workshop 2026
  8. Please Don’t Kill My Vibe: Empowering Agents with Data Flow Control
    Charlie Summers, Haneen Mohammed, Eugene Wu
    CIDR 2026 Slides
  9. A Data Aggregation Visualization System supported by Processing-in-Memory.
    Junyoung Kim, Madhulika Balakumar, Kenneth Ross
    Sixteenth International Workshop on Accelerating Analytics and Data Management Systems Using Modern Processor and Storage Architectures (ADMS at VLDB 2025)
  10. Lineage Capture Trade-offs: A Case Study in DuckDB
    Haneen Mohammed, Eugene Wu
    ProvWeek at SIGMOD 2025
  11. DEMO: Fast Hypothetical Update Evaluation
    Haneen Hohammed, Alexander Yao, Charlie Summers, Hongbin Zhong, Gramit Yeuk-Yin Chan, Subrata Mitra, Lampros Flokas, Eugene Wu
    ProvWeek at SIGMOD 2025
  12. Compression for High-Performance Lineage
    Xiaoyu Han, Haneen Mohammed, Charlie Summers, Eugene Wu
    ProvWeek at SIGMOD 2025
  13. Suna: Scalable Causal Confounder Discovery over Relational Data
    Jiaxiang Liu, Siyuan Xia, Daniel Alabi, Eugene Wu
    VLDB 2025
  14. Tutorial: Privacy and Security in Distributed Data Markets
    Daniel Alabi, Sainyam Galhotra, Shagufta Mehnaz, Zeyu Song, Eugene Wu
    SIGMOD 2025 tutorial Code
  15. FaDE: More Than a Million What-ifs Per Second
    Haneen Mohammed, Alexander Yao, Lampros Flokas, Charlie Summers, Hongbin Zhong, Gramit Chan, Surata Mitra, Eugene Wu
    VLDB 2025
  16. SET: Searching Effective Supervised Learning Augmentations in Large Tabular Data Repositories
    Jerry Liu, Zachary Huang, Eugene Wu
    GUIDEAI Workshop at SIGMOD 2024
  17. Lightweight Materialization for Fast Dashboards Over Joins
    Zachary Huang, Eugene Wu
    SIGMOD 2024
  18. The Fast and the Private: Task-based Dataset Search
    Zezhou Huang, Jiaxiang Liu, Haonan Wang, Eugene Wu
    CIDR 2024 Slides
  19. Saibot: A Differentially Private Data Search Platform
    Zezhou Huang, Jiaxiang Liu, Daniel Alabi, Raul Castro Fernandez, Eugene Wu
    VLDB 2023