This repository collects documented instances of failures exhibited by ChatGPT and related large language models. The repository aims to provide a readily accessible archive of errors, inconsistencies, and unexpected behaviors. It is intended to facilitate comparison across models and aid in the development of more robust and reliable AI systems. The core approach involves documenting specific failure cases with links to original reports.
The repository includes a wide range of failures, from logical reasoning errors and factual inaccuracies to biases and unexpected responses. It specifically focuses on capturing real-world examples from social media and online forums. The structure allows for easy browsing and comparison of different types of failures within both ChatGPT and older iterations of the model. The inclusion of links allows for reproducibility and detailed investigation of each case.
- Failure Cases: Documents specific instances where models produce incorrect, nonsensical or inappropriate responses.
- Date-Based Organization: Organizes failures chronologically to track model improvement and identify persistent issues.
- Reproducibility Information: Includes links to original posts where available, aiding in reproducing failure conditions.
- Bias Documentation: Highlights instances of harmful biases in model outputs.
- Error Types: Categorizes failures into different categories (e.g., arithmetic, reasoning, bias) for easier analysis.
The repository is actively maintained with new failures added regularly. The focus is on documenting recent issues and preserving historical examples. While not a fully formalized research project, it serves as a valuable resource for developers and researchers interested in evaluating and mitigating model limitations. Community contributions are encouraged through submissions of new failure cases.
This repository benefits developers seeking to understand the shortcomings of current LLMs and researchers looking for data to analyze model behavior. It is valuable for anyone building applications with these models, particularly in domains requiring high reliability and accuracy. By observing these failures, users can better anticipate and prepare for potential issues when deploying these technologies.
