Treffer: A case study using sewage metagenomic data for assessment of text-to-SQL capabilities in large language models.
Nat Commun. 2024 Aug 30;15(1):7551. (PMID: 39215001)
Emerg Infect Dis. 2021 May;27(5):1405-1415. (PMID: 33900177)
ISME J. 2017 Dec;11(12):2864-2868. (PMID: 28742071)
Sci Data. 2016 Mar 15;3:160018. (PMID: 26978244)
Nat Methods. 2023 Aug;20(8):1203-1212. (PMID: 37500759)
Nat Commun. 2019 Mar 8;10(1):1124. (PMID: 30850636)
Gigascience. 2019 Jun 1;8(6):. (PMID: 31220250)
Bioinformatics. 2014 Jul 15;30(14):2068-9. (PMID: 24642063)
Nat Biotechnol. 2017 Aug 8;35(8):725-731. (PMID: 28787424)
Nat Rev Microbiol. 2023 Apr;21(4):213-214. (PMID: 36470999)
Nat Commun. 2022 Dec 1;13(1):7251. (PMID: 36456547)
BMC Bioinformatics. 2018 Aug 29;19(1):307. (PMID: 30157759)
Sci Total Environ. 2023 May 15;873:162209. (PMID: 36796689)
Bioinformatics. 2019 Nov 15;:. (PMID: 31730192)
Weitere Informationen
Relational databases offer an efficient solution for storing and retrieving complex data sets, yet the requirement for SQL programming expertise presents a significant challenge for many life science users. We explore whether a cutting-edge large language model can effectively translate plain English queries into SQL scripts (Text-to-SQL), thereby simplifying database interaction and eliminating the typical usage barriers. A complex database comprising 19 interconnected tables of metagenomic analyses from 239 sewage samples across five European cities was available. A large language model was provided with details of the database's structure and background information on its contents. We evaluated the functionalities of this "SewageGPT" tool and assessed the accuracy of its responses to complex questions and visualisation of results. Providing a detailed description of the database enabled SewageGPT to accurately respond to complex inquiries, accelerating the database querying process. Knowledge of the database content proved beneficial, as it minimized the risk of ambiguities in queries; however, ambiguities can lead to incorrect responses. Therefore, human oversight remains crucial, particularly for questions that lack detail or involve ambiguities. The integration of state-of-the-art large language models with direct database connectivity substantially enhances the efficiency of query generation, statistical analysis and visualization of the results.
(© 2025. The Author(s).)
Declarations. Competing interests: The authors declare no competing interests.