bugbug/README.md

# bugbug

## Classifiers
- **bug vs feature** - Bugs on Bugzilla aren't always bugs. Sometimes they are feature requests, refactorings, and so on. The aim of this classifier is to distinguish between bugs that are actually bugs and bugs that aren't. The dataset currently contains 2110 bugs, the accuracy of the current classifier is ~93% (precision ~95%, recall ~94%).

- **defect vs feature vs task** - Extension of the previous classifier to detect differences also between feature requests and development tasks.

- **component** - The aim of this classifier is to assign product/component to (untriaged) bugs.

- **regression vs non-regression** - Bugzilla has a `regression` keyword to identify bugs that are regressions. Unfortunately it isn't used consistently. The aim of this classifier is to detect bugs that are regressions.

- **tracking** - The aim of this classifier is to detect bugs to track.

- **uplift** - The aim of this classifier is to detect bugs for which uplift should be approved and bugs for which uplift should not be approved.

- **devdocneeded** - The aim of this classifier is to detect bugs which should be documented for developers.

- **qaneeded** - The aim of this classifier is to detect bugs that would need QA verification.

- **bugtype** - The aim of this classifier is to classify bugs according to their type.

## Setup

Run `pip install -r requirements.txt` and `pip install -r test-requirements.txt`

If you update the bugs database, run `xz -v9 -k data/bugs.json`.
If you update the commits database, run `xz -v9 -k data/commits.json`.


## Usage

Run the `run.py` script to perform training / classification. The first time `run.py` is executed, the `--train` argument should be used to automatically download databases containing bugs and commits data.


### Running the repository mining script

1. Clone https://hg.mozilla.org/mozilla-central/.
2. Run `./mach vcs-setup` in the directory where you have cloned mozilla-central.
3. Enable the pushlog, hgmo and mozext extensions. For example, if you are on Linux, add the following to the extensions section of the `~/.hgrc` file:
    ```
    pushlog = ~/.mozbuild/version-control-tools/hgext/pushlog
    hgmo = ~/.mozbuild/version-control-tools/hgext/hgmo
    mozext = ~/.mozbuild/version-control-tools/hgext/mozext
    firefoxtree = ~/.mozbuild/version-control-tools/hgext/firefoxtree
    ```
3. Run the `repository.py` script, with the only argument being the path to the mozilla-central repository.

Note: the script will take a long time to run (on my laptop more than 7 hours). If you want to test a simple change and you don't intend to actually mine the data, you can modify the repository.py script to limit the number of analyzed commits. Simply add `limit=1024` to the call to the `log` command.


## Structure of the project
- `bugbug/labels` contains manually collected labels;
- `bugbug/db.py` is an implementation of a really simple JSON database;
- `bugbug/bugzilla.py` contains the functions to retrieve bugs from the Bugzilla tracking system;
- `bugbug/repository.py` contains the functions to mine data from the mozilla-central (Firefox) repository;
- `bugbug/bug_features.py` contains functions to extract features from bug/commit data;
- `bugbug/model.py` contains the base class that all models derive from;
- `bugbug/models` contains implementations of specific models;
- `bugbug/nn.py` contains utility functions to include Keras models into a scikit-learn pipeline;
- `bugbug/utils.py` contains misc utility functions;
- `bugbug/nlp` contains utility functions for NLP;
- `bugbug/labels.py` contains utility functions for handling labels;
- `bugbug/bug_snapshot.py` contains a module to play back the history of a bug.

## Auto-formatting setup

This project is using [pre-commit](https://pre-commit.com/). Please run `pre-commit install` to install the git pre-commit hooks on your clone.

Then every time you will try to commit, it will check that the files are correctly formatted before letting you commit.
Update documentation to match what the project became 2018-12-05 18:25:13 +03:00			`# bugbug`
First commit Former-commit-id: 37f2820f781ae0e6719cf7677e90c0a189f026e0 2018-03-11 23:12:35 +03:00
Update documentation to match what the project became 2018-12-05 18:25:13 +03:00			`## Classifiers`
			`- bug vs feature - Bugs on Bugzilla aren't always bugs. Sometimes they are feature requests, refactorings, and so on. The aim of this classifier is to distinguish between bugs that are actually bugs and bugs that aren't. The dataset currently contains 2110 bugs, the accuracy of the current classifier is ~93% (precision ~95%, recall ~94%).`

Add more docs about other classifiers 2019-02-20 03:10:12 +03:00			`- defect vs feature vs task - Extension of the previous classifier to detect differences also between feature requests and development tasks.`

			`- component - The aim of this classifier is to assign product/component to (untriaged) bugs.`

Update documentation to match what the project became 2018-12-05 18:25:13 +03:00			- regression vs non-regression - Bugzilla has a `regression` keyword to identify bugs that are regressions. Unfortunately it isn't used consistently. The aim of this classifier is to detect bugs that are regressions.

			`- tracking - The aim of this classifier is to detect bugs to track.`
First commit Former-commit-id: 37f2820f781ae0e6719cf7677e90c0a189f026e0 2018-03-11 23:12:35 +03:00
Add uplift model to docs 2018-12-21 16:46:42 +03:00			`- uplift - The aim of this classifier is to detect bugs for which uplift should be approved and bugs for which uplift should not be approved.`

Add more docs about other classifiers 2019-02-20 03:10:12 +03:00			`- devdocneeded - The aim of this classifier is to detect bugs which should be documented for developers.`

			`- qaneeded - The aim of this classifier is to detect bugs that would need QA verification.`

Move Contributing section to CONTRIBUTING.md (#406) Also add docs about bug type classifier to README.md 2019-05-14 16:51:31 +03:00			`- bugtype - The aim of this classifier is to classify bugs according to their type.`
Use MongoDB to store bugs Former-commit-id: c9e742744e960fe9afab609fa3c52833ebc1810b 2018-09-21 18:11:34 +03:00
			`## Setup`

Update setup steps 2018-11-20 19:24:02 +03:00			Run `pip install -r requirements.txt` and `pip install -r test-requirements.txt`
Use MongoDB to store bugs Former-commit-id: c9e742744e960fe9afab609fa3c52833ebc1810b 2018-09-21 18:11:34 +03:00
Remove data files from the repo and host them on another service Former-commit-id: db43a17445659d664fc41c014e9fe2d61c98b4ba 2018-11-20 12:36:04 +03:00			If you update the bugs database, run `xz -v9 -k data/bugs.json`.
			If you update the commits database, run `xz -v9 -k data/commits.json`.
Add 'Usage' paragraph 2018-12-22 04:34:09 +03:00

			`## Usage`

Update README.md (#70) 2019-01-16 23:46:52 +03:00			Run the `run.py` script to perform training / classification. The first time `run.py` is executed, the `--train` argument should be used to automatically download databases containing bugs and commits data.
Explain how to run the repository mining script 2019-01-30 19:56:23 +03:00

			`### Running the repository mining script`

			`1. Clone https://hg.mozilla.org/mozilla-central/.`
			2. Run `./mach vcs-setup` in the directory where you have cloned mozilla-central.
Update docs for running the repository mining script (#209) 2019-03-07 22:04:20 +03:00			3. Enable the pushlog, hgmo and mozext extensions. For example, if you are on Linux, add the following to the extensions section of the `~/.hgrc` file:
			```
Remove unnecessary indentation 2019-03-19 02:56:59 +03:00			`pushlog = ~/.mozbuild/version-control-tools/hgext/pushlog`
			`hgmo = ~/.mozbuild/version-control-tools/hgext/hgmo`
			`mozext = ~/.mozbuild/version-control-tools/hgext/mozext`
Add firefoxtree extension to list of needed Mercurial extensions Fixes #212 2019-03-19 02:58:02 +03:00			`firefoxtree = ~/.mozbuild/version-control-tools/hgext/firefoxtree`
Update docs for running the repository mining script (#209) 2019-03-07 22:04:20 +03:00			```
Explain how to run the repository mining script 2019-01-30 19:56:23 +03:00			3. Run the `repository.py` script, with the only argument being the path to the mozilla-central repository.

Update docs to mention the log command instead of hg.log (#159) 2019-02-08 18:55:30 +03:00			Note: the script will take a long time to run (on my laptop more than 7 hours). If you want to test a simple change and you don't intend to actually mine the data, you can modify the repository.py script to limit the number of analyzed commits. Simply add `limit=1024` to the call to the `log` command.
Add short description of all the files in the repository 2019-01-30 20:02:46 +03:00

			`## Structure of the project`
			- `bugbug/labels` contains manually collected labels;
			- `bugbug/db.py` is an implementation of a really simple JSON database;
			- `bugbug/bugzilla.py` contains the functions to retrieve bugs from the Bugzilla tracking system;
			- `bugbug/repository.py` contains the functions to mine data from the mozilla-central (Firefox) repository;
			- `bugbug/bug_features.py` contains functions to extract features from bug/commit data;
			- `bugbug/model.py` contains the base class that all models derive from;
			- `bugbug/models` contains implementations of specific models;
			- `bugbug/nn.py` contains utility functions to include Keras models into a scikit-learn pipeline;
			- `bugbug/utils.py` contains misc utility functions;
			- `bugbug/nlp` contains utility functions for NLP;
			- `bugbug/labels.py` contains utility functions for handling labels;
			- `bugbug/bug_snapshot.py` contains a module to play back the history of a bug.
Pre commit setup (#252) * Add pre-commit configuration Add auto-formatting configuration using the https://pre-commit.com/ project. Having auto-formatting setup and automatically enforced helps speeding up development and review process. * Apply the auto-formatting on all files in the repository * Removes flake8-quotes as it conflicts with Black formatting * Disable some Flake8 rules Disable Flake8 rules that are handled by Black. The list comes from https://github.com/ambv/black/issues/429#issuecomment-472687803. 2019-04-09 16:57:29 +03:00
			`## Auto-formatting setup`

			This project is using [pre-commit](https://pre-commit.com/). Please run `pre-commit install` to install the git pre-commit hooks on your clone.

Change good-first-bug label name in README 2019-05-13 18:40:37 +03:00			`Then every time you will try to commit, it will check that the files are correctly formatted before letting you commit.`