- Preprocessing time30 minto3 min
- Accuracy of the column-name prediction model85%to96%
Software development company (AI service for the construction industry)
Improving an AI that predicts CO2 emissions from construction estimates, and building its inference API
Background
This AI service predicts CO2 emissions from construction cost estimates (Excel). Because the estimate format differed for each tenant, preprocessing logic and models were built separately per tenant. Preprocessing was slow to run and costly to maintain. On top of that, there was no way to search construction procedure manuals, which contain figures and long tables, for the information needed.
What we did
- Rewrote preprocessing that relied heavily on loops to use vectorized operations in Pandas.
- Built a retrieval-augmented generation (RAG) pipeline for manuals with figures and tables: Amazon Bedrock inserts summaries of figures into the text, and tables are split while preserving their structure before being indexed for search.
- Consolidated scattered experiment environments on Databricks and used MLflow to record experiment parameters alongside model accuracy, so any model change can be checked for accuracy regressions.
- Identified improvements to the model architecture and preprocessing of the AI that predicts column names in estimates.
- Organized inference as a FastAPI API and handled model deployment to Google Cloud Run.
Results
Preprocessing time dropped from 30 minutes to 3 minutes. Accuracy of the column-name prediction AI improved from 85% to 96%. With experiment conditions and results now traceable, the groundwork is in place to move to analysis that no longer depends on per-tenant customization.

