Platform for Machine Learning projects on Software Engineering

hacktoberfest machine-learning software-engineering

Перейти к файлу

Marco Castelluccio d9cdcdc238 Enable feature importance calculation for the defect/enhancement/task model		2019-07-11 20:44:07 +02:00
bugbug	Enable feature importance calculation for the defect/enhancement/task model	2019-07-11 20:44:07 +02:00
http_service	Rename the suggestion field into class (#670 )	2019-07-04 12:49:33 +02:00
infra	Enable feature importance calculation for the defect/enhancement/task model	2019-07-11 20:44:07 +02:00
scripts	Store most important features in a JSON file too	2019-07-10 14:57:16 +02:00
tests	Use some bug features in the Backout model (#615 )	2019-07-11 15:42:02 +02:00
.dockerignore	Update root dockerignore (#266 )	2019-04-12 12:31:52 +02:00
.flake8	Remove useless comment	2019-06-11 17:28:45 +02:00
.gitignore	Add .coverage file to .gitignore	2019-06-11 20:06:14 +02:00
.isort.cfg	Add word2vec with WMD distance similarity (#666 )	2019-07-08 12:24:13 +02:00
.pre-commit-config.yaml	Add .taskcluster.yml validator as a pre-commit check (#567 )	2019-06-13 18:41:04 +02:00
.taskcluster.yml	Update task-boot to 0.1.9 (#675 )	2019-07-05 15:36:16 +02:00
CODE_OF_CONDUCT.md	Remove trailing whitespaces from CODE_OF_CONDUCT.md	2019-05-14 17:31:02 +02:00
CONTRIBUTING.md	Remove 'reserved-for-beginners' rule	2019-07-05 11:39:24 +02:00
LICENSE	First commit	2018-03-11 20:12:35 +00:00
MANIFEST.in	Move Keras as an optional dependency (#319 )	2019-04-26 18:50:02 +02:00
README.md	Add a short description about bugbug (#574 )	2019-06-10 11:05:30 +02:00
VERSION	Version 0.0.65	2019-07-09 21:32:36 +02:00
docker-compose.yml	Use the base image for training models (#656 )	2019-06-29 00:01:51 +02:00
extra-nlp-requirements.txt	Update nltk from 3.4.3 to 3.4.4 (#673 )	2019-07-05 12:12:43 +02:00
extra-nn-requirements.txt	Update tensorflow from 1.13.1 to 1.14.0 (#599 )	2019-06-18 22:49:19 -07:00
pytest.ini	Add a mock DB for tests to avoid downloading the full DB (#273 )	2019-04-18 14:01:25 +02:00
requirements.txt	Update imbalanced-learn from 0.4.3 to 0.5.0 (#657 )	2019-06-28 23:53:02 +02:00
run.py	For feature importance, show human readable feature names	2019-07-09 11:16:18 +02:00
setup.py	Make commit_classifier a console_script	2019-07-02 19:39:59 +02:00
test-requirements.txt	Improve triggerSchema of the hooks and test it	2019-07-03 17:26:01 +02:00

README.md

bugbug

Bugbug aims at leveraging machine learning techniques to help with bug and quality management, and other software engineering tasks.

More information on the Mozilla hacks blog: https://hacks.mozilla.org/2019/04/teaching-machines-to-triage-firefox-bugs/

Classifiers

assignee - The aim of this classifier is to suggest an appropriate assignee for a bug.
backout - The aim of this classifier is to detect patches that might be more likely to be backed-out (because of build or test failures). It could be used for test prioritization/scheduling purposes.
bugtype - The aim of this classifier is to classify bugs according to their type.
component - The aim of this classifier is to assign product/component to (untriaged) bugs.
defect vs enhancement vs task - Extension of the defect classifier to detect differences also between feature requests and development tasks.
defect - Bugs on Bugzilla aren't always bugs. Sometimes they are feature requests, refactorings, and so on. The aim of this classifier is to distinguish between bugs that are actually bugs and bugs that aren't. The dataset currently contains 2110 bugs, the accuracy of the current classifier is ~93% (precision ~95%, recall ~94%).
devdocneeded - The aim of this classifier is to detect bugs which should be documented for developers.
duplicate - The aim of this classifier is to detect duplicate bugs.
qaneeded - The aim of this classifier is to detect bugs that would need QA verification.
regression vs non-regression - Bugzilla has a regression keyword to identify bugs that are regressions. Unfortunately it isn't used consistently. The aim of this classifier is to detect bugs that are regressions.
regressionrange - The aim of this classifier is to detect regression bugs that have a regression range vs those that don't.
regressor - The aim of this classifier is to detect patches which are more likely to cause regressions. It could be used to make riskier patches undergo more scrutiny.
stepstoreproduce - The aim of this classifier is to detect bugs that have steps to reproduce vs those that don't.
tracking - The aim of this classifier is to detect bugs to track.
uplift - The aim of this classifier is to detect bugs for which uplift should be approved and bugs for which uplift should not be approved.

Setup

Run pip install -r requirements.txt and pip install -r test-requirements.txt

Auto-formatting

This project is using pre-commit. Please run pre-commit install to install the git pre-commit hooks on your clone.

Every time you will try to commit, pre-commit will run checks on your files to make sure they follow our style standards and they aren't affected by some simple issues. If the checks fail, pre-commit won't let you commit.

Usage

Run the run.py script to perform training / classification. The first time run.py is executed, the --train argument should be used to automatically download databases containing bugs and commits data.

Running the repository mining script

Clone https://hg.mozilla.org/mozilla-central/.
Run ./mach vcs-setup in the directory where you have cloned mozilla-central.

Enable the pushlog, hgmo and mozext extensions. For example, if you are on Linux, add the following to the extensions section of the ~/.hgrc file:

pushlog = ~/.mozbuild/version-control-tools/hgext/pushlog
hgmo = ~/.mozbuild/version-control-tools/hgext/hgmo
mozext = ~/.mozbuild/version-control-tools/hgext/mozext
firefoxtree = ~/.mozbuild/version-control-tools/hgext/firefoxtree

Run the repository.py script, with the only argument being the path to the mozilla-central repository.

Note: the script will take a long time to run (on my laptop more than 7 hours). If you want to test a simple change and you don't intend to actually mine the data, you can modify the repository.py script to limit the number of analyzed commits. Simply add limit=1024 to the call to the log command.

Structure of the project

bugbug/labels contains manually collected labels;
bugbug/db.py is an implementation of a really simple JSON database;
bugbug/bugzilla.py contains the functions to retrieve bugs from the Bugzilla tracking system;
bugbug/repository.py contains the functions to mine data from the mozilla-central (Firefox) repository;
bugbug/bug_features.py contains functions to extract features from bug/commit data;
bugbug/model.py contains the base class that all models derive from;
bugbug/models contains implementations of specific models;
bugbug/nn.py contains utility functions to include Keras models into a scikit-learn pipeline;
bugbug/utils.py contains misc utility functions;
bugbug/nlp contains utility functions for NLP;
bugbug/labels.py contains utility functions for handling labels;
bugbug/bug_snapshot.py contains a module to play back the history of a bug.