Data cleaning is one of the most time-consuming stages of the data lifecycle, yet it is essential for producing accurate and reliable analysis. Recent developments in large language models (LLMs) have created new opportunities to support data wrangling activities by automating or assisting tasks that previously relied heavily on manual effort. Research suggests that LLMs have the potential to improve efficiency and reduce the complexity of preparing data for analysis, particularly when combined with clear instructions and human oversight (Liu et al., 2024; Zhang et al., 2024).
AI tools can support a variety of common data cleaning activities. These include identifying missing values, detecting duplicate records, standardising inconsistent text, and suggesting transformations to improve data quality. LLMs can also generate formulas, SQL queries or code to automate repetitive tasks. Rather than replacing traditional data preparation methods, these capabilities provide additional support that can help users work more efficiently and focus more time on analysis and decision-making (Liu et al., 2024).
The quality of these outputs depends heavily on the instructions provided. Effective prompts for data cleaning typically contain four key elements. First, context helps the model understand the type of dataset being analysed. Second, the task should clearly state the objective, such as identifying duplicate customer records or standardising date formats. Third, any rules or constraints should be specified to ensure recommendations align with business requirements. Finally, defining the desired output format helps present results in a clear and usable way. Research into prompt engineering highlights that well-structured instructions improve the relevance and consistency of AI-generated outputs (Chen et al., 2025; Sahoo et al., 2024).
By providing context, constraints and an expected output format, the prompt gives the AI model clear guidance and helps reduce ambiguity. This increases the likelihood of receiving useful recommendations rather than generic responses (Chen et al., 2025).
Despite these benefits, AI-generated outputs should be treated as recommendations rather than automatic corrections. LLMs may misunderstand context or suggest changes that do not align with organisational requirements. Research into automated data wrangling emphasises the importance of keeping humans involved in reviewing and validating outputs before applying them to datasets (Liu et al., 2024; Zhang et al., 2024). Users remain responsible for ensuring that changes improve data quality and do not introduce additional errors.
Prompt engineering for data cleaning therefore represents a practical collaboration between humans and AI. By combining domain knowledge with clear instructions, users can use AI tools to support repetitive tasks, increase efficiency and improve consistency. However, the quality of the final dataset still depends on critical thinking, validation and informed decision-making rather than automation alone.
Action Point
Choose a dataset you regularly work with and identify one data quality issue, such as missing values, duplicate records or inconsistent formats. Create a prompt that provides context about the dataset, clearly defines the cleaning task and specifies any rules or constraints. Use an AI tool to generate recommendations, then review and validate the output before deciding which changes should be applied. Reflect on how AI supported the process and where human judgement was still required.