FLAML

Граф коммитов

Автор	SHA1	Сообщение	Дата
Daniel Grindrod	5a74227bc3	Flaml: fix lgbm reproducibility (#1369 ) * fix: Fixed bug where every underlying LGBMRegressor or LGBMClassifier had n_estimators = 1 * test: Added test showing case where FLAMLised CatBoostModel result isn't reproducible * fix: Fixing issue where callbacks cause LGBM results to not be reproducible * Update test/automl/test_regression.py Co-authored-by: Li Jiang <bnujli@gmail.com> * fix: Adding back the LGBM EarlyStopping * refactor: Fix tweaked to ensure other models aren't likely to be affected * test: Fixed test to allow reproduced results to be better than the FLAML results, when LGBM earlystopping is involved --------- Co-authored-by: Daniel Grindrod <Daniel.Grindrod@evotec.com> Co-authored-by: Li Jiang <bnujli@gmail.com>	2024-11-01 10:06:15 +08:00
Ranuga	7644958e21	Add documentation for `automl.model.estimator` usage (#1311 ) * Added documentation for automl.model.estimator usage Updated documentation across various examples and the model.py file to include information about automl.model.estimator. This addition enhances the clarity and usability of FLAML by providing users with clear guidance on how to utilize this feature in their AutoML workflows. These changes aim to improve the overall user experience and facilitate easier understanding of FLAML's capabilities. * fix: Ran pre-commit hook on docs --------- Co-authored-by: Li Jiang <bnujli@gmail.com> Co-authored-by: Daniel Grindrod <dannycg1996@gmail.com> Co-authored-by: Daniel Grindrod <Daniel.Grindrod@evotec.com>	2024-10-31 20:53:54 +08:00
Daniel Grindrod	a316f84fe1	fix: LinearSVC results now reproducible (#1376 ) Co-authored-by: Daniel Grindrod <Daniel.Grindrod@evotec.com>	2024-10-31 14:02:16 +08:00
Daniel Grindrod	72881d3a2b	fix: Fixing the random state of ElasticNetClassifier by default, to ensure reproduciblity. Also included elasticnet in reproducibility tests (#1374 ) Co-authored-by: Daniel Grindrod <Daniel.Grindrod@evotec.com> Co-authored-by: Li Jiang <bnujli@gmail.com>	2024-10-29 14:21:43 +08:00
Li Jiang	69da685d1e	Fix data transform issue, spark log_loss metric compute error and json dumps TypeError (Sync Fabric till 3c545e67) (#1371 ) * Merged PR 1444697: Fix json dumps TypeError Fix json dumps TypeError ---- Bug fix to address a `TypeError` in `json.dumps`. This pull request fixes a `TypeError` encountered when using `json.dumps` on `automl._automl_user_configurations` by introducing a safe JSON serialization function. - Added `safe_json_dumps` function in `flaml/fabric/mlflow.py` to handle non-serializable objects. - Updated `MLflowIntegration` class in `flaml/fabric/mlflow.py` to use `safe_json_dumps` for JSON serialization. - Modified `test/automl/test_multiclass.py` to test the new `safe_json_dumps` function. Related work items: #3439408 * Fix data transform issue and spark log_loss metric compute error	2024-10-29 11:58:40 +08:00
Li Jiang	9724c626cc	Remove outdated comment (#1366 )	2024-10-24 12:17:21 +08:00
Daniel Grindrod	d224218ecf	fix: FLAML catboost metrics arent reproducible (#1364 ) * fix: CatBoostRegressors metrics are now reproducible * test: Made tests live, which ensure the reproducibility of catboost models * fix: Added defunct line of code as a comment * fix: Re-adding removed if statement, and test to show one issue that if statement can cause * fix: Stopped ending CatBoost training early when time budget is running out --------- Co-authored-by: Daniel Grindrod <Daniel.Grindrod@evotec.com>	2024-10-23 13:51:23 +08:00
Daniel Grindrod	5c0f18b7bc	fix: Cross validation process isn't always run to completion (#1360 )	2024-10-01 08:24:53 +08:00
Li Jiang	49ba962d47	Support logger_formatter without automl dependencies (#1356 )	2024-09-21 20:04:46 +08:00
Li Jiang	5bfa0b1cd3	Improve mlflow integration and add more models (#1331 ) * Add more spark models and improved mlflow integration * Update test_extra_models, setup and gitignore * Remove autofe * Remove autofe * Remove autofe * Sync changes in internal * Fix test for env without pyspark * Fix import errors * Fix tests * Fix typos * Fix pytorch-forecasting version * Remove internal funcs, rename _mlflow.py * Fix import error * Fix dependency * Fix experiment name setting * Fix dependency * Update pandas version * Update pytorch-forecasting version * Add warning message for not has_automl * Fix test errors with nltk 3.8.2 * Don't enable mlflow logging w/o an active run * Fix pytorch-forecasting can't be pickled issue * Update pyspark tests condition * Update synapseml * Update synapseml * No parent run, no logging for OSS * Log when autolog is enabled * upgrade code * Enable autolog for tune * Increase time budget for test * End run before start a new run * Update parent run * Fix import error * clean up * skip macos and win * Update notes * Update default value of model_history	2024-08-13 07:53:47 +00:00
Jirka Borovec	b348cb1136	configure & apply pyupgrade with `py3.8+` (#1333 ) * configure pyupgrade with `py3.8+` * apply update --------- Co-authored-by: Li Jiang <bnujli@gmail.com>	2024-08-12 02:54:18 +00:00
Yang, Bo	853c9501bc	Keep searching hyperparameters when `r2_score` raises an error (#1325 ) * Keep searching hyperparameters when `r2_score` raises an error * Add log info --------- Co-authored-by: Li Jiang <bnujli@gmail.com>	2024-08-06 15:01:10 +00:00
Yang, Bo	8e63dd417b	Don't pass `callbacks=None` to `XGBoostSklearnEstimator._fit` (#1322 ) * Don't pass `callbacks=None` to `XGBoostSklearnEstimator._fit` The original implmentation would pass `callbacks=None` to `XGBoostSklearnEstimator._fit` and eventually lead to a `TypeError` of `XGBModel.fit() got an unexpected keyword argument 'callbacks'`. This PR instead does not pass the `callbacks=None` parameter to avoid the error. * Update setup.py to allow for xgboost 2.x --------- Co-authored-by: Li Jiang <bnujli@gmail.com>	2024-08-06 09:24:11 +00:00
Li Jiang	a68d073ccf	Add support to python 3.11 (#1326 ) * Add support to python 3.11 * Fix workflow python version comparison * Ray is not supported in python 3.11 * Fix test_numpy	2024-07-31 00:18:41 +00:00
Ranuga	67f4048667	Update ts_model.py (#1312 ) Co-authored-by: Li Jiang <bnujli@gmail.com>	2024-07-22 05:32:51 +00:00
Li Jiang	d8129b9211	Fix typos, upgrade yarn packages, add some improvements (#1290 ) * Fix typos, upgrade yarn packages, add some improvements * Fix joblib 1.4.0 breaks joblib-spark * Fix xgboost test error * Pin xgboost<2.0.0 * Try update prophet to 1.5.1 * Update github workflow * Revert prophet version * Update github workflow * Update install libomp * Fix test errors * Fix test errors * Add retry to test and coverage * Revert "Add retry to test and coverage" This reverts commit `ce13097cd5`. * Increase test budget * Add more data to test_models, try fixing ValueError: Found array with 0 sample(s) (shape=(0, 252)) while a minimum of 1 is required.	2024-07-19 13:40:04 +00:00
Jirka Borovec	165d7467f9	precommit: introduce `mdformat` (#1276 ) * precommit: introduce `mdformat` * precommit: apply	2024-03-19 22:46:56 +00:00
Gleb Levitski	3de0dc667e	Add ruff sort to pre-commit and sort imports in the library (#1259 ) * lint * bump ver * bump ver * fixed circular import --------- Co-authored-by: Jirka Borovec <6035284+Borda@users.noreply.github.com>	2024-03-12 21:28:57 +00:00
Chi Wang	1a9fa3ac23	Np.inf (#1289 ) * np.Inf -> np.inf * bump version to 2.1.2	2024-03-12 16:27:05 +00:00
Gleb Levitski	6b93c2e394	[ENH] Add support for sklearn HistGradientBoostingEstimator (#1230 ) * Update model.py HistGradientBoosting support * Create __init__.py * Update model.py * Create histgb.py * Update __init__.py * Update test_model.py * added histgb to estimator list * Update Task-Oriented-AutoML.md added docs * lint * fixed bugs --------- Co-authored-by: Gleb <gleb@Glebs-MacBook-Pro.local> Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-10-31 14:45:23 +00:00
Chi Wang	fda9fa0103	improve docstr of preprocessors (#1227 ) * improve docstr of preprocessors * Update SynapseML version * RFix test --------- Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-09-29 03:07:21 +00:00
Dominik Moritz	46162578f8	Fix typo Whetehr -> Whether (#1220 ) Co-authored-by: Chi Wang <wang.chi@microsoft.com>	2023-09-22 15:27:02 +00:00
Chi Wang	868e7dd1ca	support xgboost 2.0 (#1219 ) * support xgboost 2.0 * try classes_ * test version * quote * use_label_encoder * Fix xgboost test error * remove deprecated files * remove deprecated files * remove deprecated import * replace deprecated import in integrate_spark.ipynb * replace deprecated import in automl_lightgbm.ipynb * formatted integrate_spark.ipynb * replace deprecated import * try fix driver python path * Update python-package.yml * replace deprecated reference * move spark python env var to other section * Update setup.py, install xgb<2 for MacOS * Fix typo * assert * Try assert xgboost version * Fail fast * Keep all test/spark to try fail fast * No need to skip spark test in Mac or Win * Remove assert xgb version * Remove fail fast * Found root cause, fix test_sparse_matrix_xgboost * Revert "No need to skip spark test in Mac or Win" This reverts commit `a09034817f`. * remove assertion --------- Co-authored-by: Li Jiang <bnujli@gmail.com> Co-authored-by: levscaut <57213911+levscaut@users.noreply.github.com> Co-authored-by: levscaut <lwd2010530@qq.com> Co-authored-by: Li Jiang <lijiang1@microsoft.com>	2023-09-22 06:55:00 +00:00
Chi Wang	7ab4d114d7	silent; code_execution_config; exit; version (#1179 ) * silent; code_execution_config; exit; version * url * url * readme * preview * doc * url * endpoints * timeout * chess * Fix retrieve chat * config * mathchat --------- Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-08-14 07:09:45 +00:00
minghao	da92238ffe	Commenting use_label_encoder - xgboost (#1122 ) * Commenting use_label_encoder - xgboost * format change * moving the import xgboost version to the head * Shfit params for use_label to outside maxdept * Keep the original logic --------- Co-authored-by: Shaokun <shaokunzhang529@gmail.com> Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-08-01 01:58:30 +00:00
Chi Wang	7ddb171cd9	autogen.agent -> autogen.agentchat (#1148 ) * autogen.agent -> autogen.agentchat * bug fix in portfolio * notebook * timeout * timeout * infer lang; close #1150 * timeout * message context * context handling * add sender to generate_reply * clean up the receive function * move mathchat to contrib * contrib * last_message	2023-07-29 04:17:51 +00:00
Chi Wang	3e7aac6e8b	unify auto_reply; bug fix in UserProxyAgent; reorg agent hierarchy (#1142 ) * simplify the initiation of chat * version update * include openai * completion * load config list from json * initiate_chat * oai config list * oai config list * config list * config_list * raise_error * retry_time * raise condition * oai config list * catch file not found * catch openml error * handle openml error * handle openml error * handle openml error * handle openml error * handle openml error * handle openml error * close #1139 * use property * termination msg * AIUserProxyAgent * smaller dev container * update notebooks * match * document code execution and AIUserProxyAgent * gpt 3.5 config list * rate limit * variable visibility * remove unnecessary import * quote * notebook comments * remove mathchat from init import * two users * import location * expose config * return str not tuple * rate limit * ipython user proxy * message * None result * rate limit * rate limit * rate limit * rate limit * make auto_reply a common method for all agents * abs path * refactor and doc * set mathchat_termination * code format * modified * emove import * code quality * sender -> messages * system message * clean agent hierarchy * dict check * invalid oai msg * return * openml error * docstr --------- Co-authored-by: kevin666aa <yrwu000627@gmail.com>	2023-07-25 23:46:11 +00:00
Xiaobo Xia	eae65ac22b	suppress printing data split type (#1126 ) * first commit * second update * third update --------- Co-authored-by: “xiaoboxia” <“xiaoboxia.uni@gmail.com”> Co-authored-by: Shaokun <shaokunzhang529@gmail.com> Co-authored-by: “skzhang1” <“shaokunzhang529@gmail.com”>	2023-07-17 13:02:10 +00:00
Li Jiang	8ac9a393b8	Add log metric (#1125 ) * Add original metric to mlflow logging * Update metric	2023-07-13 12:54:39 +00:00
levscaut	5eece5c748	Enhance Integration with Spark (#1097 ) * add doc for spark * labelCol equals to label by default * change title and reformat * reference about default index type * fix doc build * Update website/docs/Examples/Integrate - Spark.md * update doc * Added more references * remove exception case when `y_train.name` is None * fix broken link --------- Co-authored-by: Wendong Li <v-wendongli@microsoft.com> Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-07-10 04:44:01 +00:00
EgorKraevTransferwise	5245efbd2c	Factor out time series-related functionality into a time series Task object (#989 ) * Refactor into automl subpackage Moved some of the packages into an automl subpackage to tidy before the task-based refactor. This is in response to discussions with the group and a comment on the first task-based PR. Only changes here are moving subpackages and modules into the new automl, fixing imports to work with this structure and fixing some dependencies in setup.py. * Fix doc building post automl subpackage refactor * Fix broken links in website post automl subpackage refactor * Fix broken links in website post automl subpackage refactor * Remove vw from test deps as this is breaking the build * Move default back to the top-level I'd moved this to automl as that's where it's used internally, but had missed that this is actually part of the public interface so makes sense to live where it was. * Re-add top level modules with deprecation warnings flaml.data, flaml.ml and flaml.model are re-added to the top level, being re-exported from flaml.automl for backwards compatability. Adding a deprecation warning so that we can have a planned removal later. * Fix model.py line-endings * WIP * WIP - Notes below Got to the point where the methods from AutoML are pulled to GenericTask. Started removing private markers and removing the passing of automl to these methods. Done with decide_split_type, started on prepare_data. Need to do the others after * Re-add generic_task * Most of the merge done, test_forecast_automl fit succeeds, fails at predict() * Remaining fixes - test_forecast.py passes * Comment out holidays-related code as it's not currently used * Further holidays cleanup * Fix imports in a test * tidy up validate_data in time series task * Test fixes * Fix tests: add Task.__str__ * Fix tests: test for ray.ObjectRef * Hotwire TS_Sklearn wrapper to fix test fail * Attempt at test fix * Fix test where val_pred_y is a list * Attempt to fix remaining tests * Push to retrigger tests * Push to retrigger tests * Push to retrigger tests * Push to retrigger tests * Remove plots from automl/test_forecast * Remove unused data size field from Task * Fix import for CLASSIFICATION in notebook * Monkey patch TFT to avoid plotting, to fix tests on MacOS * Monkey patch TFT to avoid plotting v2, to fix tests on MacOS * Monkey patch TFT to avoid plotting v2, to fix tests on MacOS * Fix circular import * remove redundant code in task.py post-merge * Fix test: set svd_solver="full" in PCA * Update flaml/automl/data.py Co-authored-by: Chi Wang <wang.chi@microsoft.com> * Fix review comments * Fix task -> str in custom learner constructor * Remove unused CLASSIFICATION imports * Hotwire TS_Sklearn wrapper to fix test fail by setting optimizer_for_horizon == False * Revert changes to the automl_classification and pin FLAML version * Fix imports in reverted notebook * Fix FLAML version in automl notebooks * Fix ml.py line endings * Fix CLASSIFICATION task import in automl_classification notebook * Uncomment pip install in notebook and revert import Not convinced this will work because of installing an older version of the package into the environment in which we're running the tests, but let's see. * Revert `c6a5dd1a0` * Fix get_classification_objective import in suggest.py * Remove hcrystallball docs reference in TS_Sklearn * Merge markharley:extract-task-class-from-automl into this * Fix import, remove smooth.py * Fix dependencies to fix TFT fail on Windows Python 3.8 and 3.9 * Add tensorboardX dependency to fix TFT fail on Windows Python 3.8 and 3.9 * Set pytorch-lightning==1.9.0 to fix TFT fail on Windows Python 3.8 and 3.9 * Set pytorch-lightning==1.9.0 to fix TFT fail on Windows Python 3.8 and 3.9 * Disable PCA reduction of lagged features for now, to fix svd convervence fail * Merge flaml/main into time_series_task * Attempt to fix formatting * Attempt to fix formatting * tentatively implement holt-winters-no covariates * fix forecast method, clean class * checking external regressors too * update test forecast * remove duplicated test file, re-add sarimax, search space cleanup * Update flaml/automl/model.py removed links. Most important one probably was: https://robjhyndman.com/hyndsight/ets-regressors/ Co-authored-by: Chi Wang <wang.chi@microsoft.com> * prevent short series * add docs * First attempt at merging Holt-Winters * Linter fix * Add holt-winters to TimeSeriesTask.estimators * Fix spark test fail * Attempt to fix another spark test fail * Attempt to fix another spark test fail * Change Black max line length to 127 * Change Black max line length to 120 * Add logging for ARIMA params, clean up time series models inheritance * Add more logging for missing ARIMA params * Remove a meaningless test causing a fail, add stricter check on ARIMA params * Fix a bug in HoltWinters * A pointless change to hopefully trigger the on and off KeyError in ARIMA.fit() * Fix formatting * Attempt to fix formatting * Attempt to fix formatting * Attempt to fix formatting * Attempt to fix formatting * Add type annotations to _train_with_config() in state.py * Add type annotations to prepare_sample_train_data() in state.py * Add docstring for time_col argument of AutoML.fit() * Address @sonichi's comments on PR * Fix formatting * Fix formatting * Reduce test time budget * Reduce test time budget * Increase time budget for the test to pass * Remove redundant imports * Remove more redundant imports * Minor fixes of points raised by Qingyun * Try to fix pandas import fail * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Try to fix pandas import fail, again * Formatting fixes * More formatting fixes * Added test that loops over TS models to ensure coverage * Fix formatting issues * Fix more formatting issues * Fix random fail in check * Put back in tests for ARIMA predict without fit * Put back in tests for lgbm * Update test/test_model.py cover dedup * Match target length to X length in missing test --------- Co-authored-by: Mark Harley <mark.harley@transferwise.com> Co-authored-by: Mark Harley <mharley.code@gmail.com> Co-authored-by: Qingyun Wu <qingyun.wu@psu.edu> Co-authored-by: Chi Wang <wang.chi@microsoft.com> Co-authored-by: Andrea W <a.ruggerini@ammagamma.com> Co-authored-by: Andrea Ruggerini <nescio.adv@gmail.com> Co-authored-by: Egor Kraev <Egor.Kraev@tw.com> Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-06-19 11:20:32 +00:00
Li Jiang	d36b2afe7f	suppress warning message of pandas_on_spark to_spark (#1058 )	2023-06-01 16:04:01 +00:00
Chi Wang	a0b318b12e	create an automl option to remove unnecessary dependency for autogen and tune (#1007 ) * version update post release v1.2.2 * automl option * import pandas * remove automl.utils * default * test * type hint and version update * dependency update * link to open in colab * use packging.version to close #725 --------- Co-authored-by: Li Jiang <lijiang1@microsoft.com> Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-05-24 23:55:04 +00:00
Susan Xueqing Liu	00c30a398e	fix NLP zero division error (#1009 ) * fix NLP zero division error * set predictions to None * set predictions to None * set predictions to None * refactor * refactor --------- Co-authored-by: Li Jiang <lijiang1@microsoft.com> Co-authored-by: Chi Wang <wang.chi@microsoft.com> Co-authored-by: Li Jiang <bnujli@gmail.com>	2023-05-03 05:50:28 +00:00
garar	31864d2d77	Add mlflow_logging param (#1015 ) Co-authored-by: Chi Wang <wang.chi@microsoft.com>	2023-05-03 03:09:04 +00:00
Jirka Borovec	73bb6e7667	pyproject.toml & switch to Ruff (#976 ) * unify config to pyproject.toml replace flake8 with Ruff * drop configs * update * fixing * Apply suggestions from code review Co-authored-by: Zvi Baratz <z.baratz@gmail.com> * setup * ci * pr template * reword --------- Co-authored-by: Zvi Baratz <z.baratz@gmail.com> Co-authored-by: Li Jiang <lijiang1@microsoft.com>	2023-04-28 01:54:55 +00:00
Susan Xueqing Liu	7114b8f742	fix zerodivision (#1000 ) * fix zerodivision * update * remove final --------- Co-authored-by: Li Jiang <lijiang1@microsoft.com>	2023-04-23 03:55:51 +00:00
Jane Illarionova	b235fe0098	Expose feature and label transformer in automl.py (#993 ) * expose label and feature transformer * linter apply * avoid undefined attribute in flaml/automl/automl.py Co-authored-by: Chi Wang <wang.chi@microsoft.com> * avoid undefined attribute in flaml/automl/automl.py Co-authored-by: Chi Wang <wang.chi@microsoft.com> * retrigger checks * retrigger checks --------- Co-authored-by: Chi Wang <wang.chi@microsoft.com>	2023-04-15 19:06:47 +00:00
Jirka Borovec	a701cd82f8	set black with 120 line length (#975 ) * set black with 120 line length * apply pre-commit * apply black	2023-04-10 19:50:40 +00:00
Susan Xueqing Liu	ef5a17cd83	handling nlp divide by zero (#926 ) * handling nlp divide by zero * catching zerodivisionerror * catching zerodivisionerror * catching zerodivisionerror * addressing comments * addressing comments * updating test case * update * add blank to last line * update nlp notebook * rerun * rerun * sync with main * add model selection for nlg * addressing keyerror * add raise exception * update * fix bug * revert * updating automl_nlp * Update flaml/automl/model.py Co-authored-by: Zvi Baratz <z.baratz@gmail.com> * address comments * address comments --------- Co-authored-by: Li Jiang <lijiang1@microsoft.com> Co-authored-by: Zvi Baratz <z.baratz@gmail.com>	2023-04-09 16:53:30 +00:00
Andrea Ruggerini	7f9402b8fd	Add Holt-Winters exponential smoothing (#962 ) * tentatively implement holt-winters-no covariates * fix forecast method, clean class * checking external regressors too * update test forecast * remove duplicated test file, re-add sarimax, search space cleanup * Update flaml/automl/model.py removed links. Most important one probably was: https://robjhyndman.com/hyndsight/ets-regressors/ Co-authored-by: Chi Wang <wang.chi@microsoft.com> * prevent short series * add docs --------- Co-authored-by: Andrea W <a.ruggerini@ammagamma.com> Co-authored-by: Chi Wang <wang.chi@microsoft.com>	2023-04-04 17:29:54 +00:00
Li Jiang	50334f2c52	Support spark dataframe as input dataset and spark models as estimators (#934 ) * add basic support to Spark dataframe add support to SynapseML LightGBM model update to pyspark>=3.2.0 to leverage pandas_on_Spark API * clean code, add TODOs * add sample_train_data for pyspark.pandas dataframe, fix bugs * improve some functions, fix bugs * fix dict change size during iteration * update model predict * update LightGBM model, update test * update SynapseML LightGBM params * update synapseML and tests * update TODOs * Added support to roc_auc for spark models * Added support to score of spark estimator * Added test for automl score of spark estimator * Added cv support to pyspark.pandas dataframe * Update test, fix bugs * Added tests * Updated docs, tests, added a notebook * Fix bugs in non-spark env * Fix bugs and improve tests * Fix uninstall pyspark * Fix tests error * Fix java.lang.OutOfMemoryError: Java heap space * Fix test_performance * Update test_sparkml to test_0sparkml to use the expected spark conf * Remove unnecessary widgets in notebook * Fix iloc java.lang.StackOverflowError * fix pre-commit * Added params check for spark dataframes * Refactor code for train_test_split to a function * Update train_test_split_pyspark * Refactor if-else, remove unnecessary code * Remove y from predict, remove mem control from n_iter compute * Update workflow * Improve _split_pyspark * Fix test failure of too short training time * Fix typos, improve docstrings * Fix index errors of pandas_on_spark, add spark loss metric * Fix typo of ndcgAtK * Update NDCG metrics and tests * Remove unuseful logger * Use cache and count to ensure consistent indexes * refactor for merge maain * fix errors of refactor * Updated SparkLightGBMEstimator and cache * Updated config2params * Remove unused import * Fix unknown parameters * Update default_estimator_list * Add unit tests for spark metrics	2023-03-25 19:59:46 +00:00
Susan Xueqing Liu	a3e770eac5	fix delete (#950 )	2023-03-14 03:19:58 +00:00
Mark Harley	27b2712016	Extract task class from automl (#857 ) * Refactor into automl subpackage Moved some of the packages into an automl subpackage to tidy before the task-based refactor. This is in response to discussions with the group and a comment on the first task-based PR. Only changes here are moving subpackages and modules into the new automl, fixing imports to work with this structure and fixing some dependencies in setup.py. * Fix doc building post automl subpackage refactor * Fix broken links in website post automl subpackage refactor * Fix broken links in website post automl subpackage refactor * Remove vw from test deps as this is breaking the build * Move default back to the top-level I'd moved this to automl as that's where it's used internally, but had missed that this is actually part of the public interface so makes sense to live where it was. * Re-add top level modules with deprecation warnings flaml.data, flaml.ml and flaml.model are re-added to the top level, being re-exported from flaml.automl for backwards compatability. Adding a deprecation warning so that we can have a planned removal later. * Fix model.py line-endings * WIP * WIP - Notes below Got to the point where the methods from AutoML are pulled to GenericTask. Started removing private markers and removing the passing of automl to these methods. Done with decide_split_type, started on prepare_data. Need to do the others after * Re-add generic_task * Fix tests: add Task.__str__ * Fix tests: test for ray.ObjectRef * Hotwire TS_Sklearn wrapper to fix test fail * Remove unused data size field from Task * Fix import for CLASSIFICATION in notebook * Update flaml/automl/data.py Co-authored-by: Chi Wang <wang.chi@microsoft.com> * Fix review comments * Fix task -> str in custom learner constructor * Remove unused CLASSIFICATION imports * Hotwire TS_Sklearn wrapper to fix test fail by setting optimizer_for_horizon == False * Revert changes to the automl_classification and pin FLAML version * Fix imports in reverted notebook * Fix FLAML version in automl notebooks * Fix ml.py line endings * Fix CLASSIFICATION task import in automl_classification notebook * Uncomment pip install in notebook and revert import Not convinced this will work because of installing an older version of the package into the environment in which we're running the tests, but let's see. * Revert `c6a5dd1a0` * Revert "Revert c6a5dd1a0" This reverts commit `e55e35adea`. * Black format model.py * Bump version to 1.1.2 in automl_xgboost * Add docstrings to the Task ABC * Fix import in custom_learner * fix 'optimize_for_horizon' for ts_sklearn * remove debugging print statements * Check for is_forecast() before is_classification() in decide_split_type * Attempt to fix formatting fail * Another attempt to fix formatting fail * And another attempt to fix formatting fail * Add type annotations for task arg in signatures and docstrings * Fix formatting * Fix linting --------- Co-authored-by: Qingyun Wu <qingyun.wu@psu.edu> Co-authored-by: EgorKraevTransferwise <egor.kraev@transferwise.com> Co-authored-by: Chi Wang <wang.chi@microsoft.com> Co-authored-by: Kevin Chen <chenkevin.8787@gmail.com>	2023-03-11 02:39:08 +00:00
Chi Wang	1ec77b58b4	improve max_valid_n and doc (#933 ) * improve max_valid_n and doc * Update README.md Co-authored-by: Li Jiang <lijiang1@microsoft.com> * newline at end of file * doc --------- Co-authored-by: Li Jiang <lijiang1@microsoft.com> Co-authored-by: Susan Xueqing Liu <liususan091219@users.noreply.github.com> Co-authored-by: Qingyun Wu <qingyun.wu@psu.edu>	2023-03-05 16:40:57 +00:00
Jirka Borovec	2ff1035733	precommit: end-of-file-fixer (#929 ) * precommit: end-of-file-fixer * exclude .gitignore * apply --------- Co-authored-by: Shaokun <shaokunzhang529@gmail.com>	2023-02-28 16:27:14 +00:00
levscaut	c6a2440348	add PySparkOvertimeMonitor to avoid exceeding time budget (#923 ) * merging * clean commit * Delete mylearner.py This file is not needed. * fix py4j import error * more tolerant cancelling time * fix problems following suggestions * Update flaml/tune/spark/utils.py Co-authored-by: Li Jiang <bnujli@gmail.com> * remove redundant model * Update test/spark/custom_mylearner.py Co-authored-by: Chi Wang <wang.chi@microsoft.com> * add docstr * reverse change in gitignore * Update test/spark/custom_mylearner.py Co-authored-by: Chi Wang <wang.chi@microsoft.com> --------- Co-authored-by: Li Jiang <bnujli@gmail.com> Co-authored-by: Chi Wang <wang.chi@microsoft.com>	2023-02-24 08:07:00 +00:00
Li Jiang	7c0340fde6	Updated dict type args default value to None (#927 )	2023-02-23 05:23:24 +00:00
Andrea Ruggerini	8e447562c7	Improve annotations in automl and ml modules (#919 ) * begin annotation in automl.py and ml.py * EstimatorSubclass + annotate metric * review: fixes + setting fit_kwargs as proper Optional * import from flaml.automl.model (import from flaml.model is deprecated) * comment n_jobs in train_estimator as well * better annotation in _compute_with_config_base Co-authored-by: Qingyun Wu <qingyun.wu@psu.edu> --------- Co-authored-by: Andrea W <a.ruggerini@ammagamma.com> Co-authored-by: Qingyun Wu <qingyun.wu@psu.edu>	2023-02-22 02:49:56 +00:00
Jirka Borovec	6aa1d16ebc	pre-commit: update config (#925 ) * update config * apply precommit	2023-02-22 00:49:38 +00:00

1 2

59 Коммитов