superbenchmark

Граф коммитов

Автор	SHA1	Сообщение	Дата
guoshzhao	f38a9829d0	ModelBenchmarks - Fix early stop logic due to num_steps. (#522 ) Description Model benchmarks can stop due to `num_steps` or `duration` config which will take effect when the value is set greater than 0. If both are set greater than 0, the earliest condition reached will work.	2023-04-28 13:15:47 +08:00
Yifan Xiong	51761b3af1	Release - SuperBench v0.8.0 (#517 ) Description Cherry-pick bug fixes from v0.8.0 to main. Major Revisions * Monitor - Fix the cgroup version checking logic (#502) * Benchmark - Fix matrix size overflow issue in cuBLASLt GEMM (#503) * Fix wrong torch usage in communication wrapper for Distributed Inference Benchmark (#505) * Analyzer: Fix bug in python3.8 due to pandas api change (#504) * Bug - Fix bug to get metric from cmd when error happens (#506) * Monitor - Collect realtime GPU power when benchmarking (#507) * Add num_workers argument in model benchmark (#511) * Remove unreachable condition when write host list (#512) * Update cuda11.8 image to cuda12.1 based on nvcr23.03 (#513) * Doc - Fix wrong unit of cpu-memory-bw-latency in doc (#515) * Docs - Upgrade version and release note (#508) Co-authored-by: guoshzhao <guzhao@microsoft.com> Co-authored-by: Ziyue Yang <ziyyang@microsoft.com> Co-authored-by: Yuting Jiang <yutingjiang@microsoft.com>	2023-04-14 12:57:55 +00:00
Ziyue Yang	8daef211dd	Benchmarks - Add distributed inference benchmark (#493 ) Description This PR adds a micro-benchmark of distributed model inference workloads. Major Revision - Add a new micro-benchmark dist-inference. - Add corresponding example and unit tests. - Update configuration files to include this new micro-benchmark. - Update micro-benchmark README. --------- Co-authored-by: Peng Cheng <chengpeng5555@outlook.com>	2023-03-24 17:15:17 +08:00
guoshzhao	a9b45a072e	Monitor - Support cgroup V2 when read system metrics. (#491 ) Description Since ubuntu 22.04 will use cgroup V2 and the file structure changed. Modify the monitor to adapt to cgroup v1 and v2.	2023-03-22 08:33:18 +00:00
Yifan Xiong	dbeba8056b	Benchmark - Support batch/shape range in cublaslt gemm (#494 ) Support batch and shape range with multiplication factors in cublaslt gemm benchmark.	2023-03-22 13:22:36 +08:00
rafsalas19	655bd0aa59	Adding HPL benchmark (#482 ) Description - Adding HPL benchmark --------- Co-authored-by: Ubuntu <azureuser@sbtestvm.jzlku1oskncengjiado35wf1hd.ax.internal.cloudapp.net> Co-authored-by: Peng Cheng <chengpeng5555@outlook.com>	2023-03-21 16:44:08 +00:00
rafsalas19	32896ca477	Adding Stream Benchmark (#473 ) Description - Added stream benchmark - Added stream unit test - Added stream example - Modified docker files to build stream --------- Co-authored-by: Ubuntu <azureuser@sbtestvm.jzlku1oskncengjiado35wf1hd.ax.internal.cloudapp.net> Co-authored-by: Peng Cheng <chengpeng5555@outlook.com> Co-authored-by: Yifan Xiong <xiongyf@yandex.com>	2023-02-13 15:34:37 -05:00
Yifan Xiong	b07fda155e	Release - SuperBench v0.7.0 (#468 ) Description Cherry-pick bug fixes from v0.7.0 to main. Major Revisions * Benchmarks - Fix missing include in FP8 benchmark (#460) * Fix bug in TE BERT model (#461) * Doc - Update benchmark doc (#465) * Bug: Fix bug for incorrect datatype judgement in cublas-function source code (#464) * Support `sb deploy` without pulling image (#466) * Docs - Upgrade version and release note (#467) Co-authored-by: Russell J. Hewett <russell.j.hewett@gmail.com> Co-authored-by: Yuting Jiang <yutingjiang@microsoft.com>	2023-01-28 11:07:06 +08:00
Yang Wang	ccccd988df	Benchmarks - Support topo-aware, pair-wise, and K-batch pattern in nccl-bw benchmark (#454 ) Support traffic patterns under the different devices in NCCL/RCCL test * change the metrics format if specified the pattern	2023-01-04 12:30:32 +00:00
Yang Wang	8e748d5649	Runner - Generate host groups file in mpi mode (#458 ) Major Revision - Add an option for pattern to generate mpi_pattern.txt file if specified the path. - In mpi pattern, serial_index and parallel_index will add in each benchmark as environment variables. Minor Revision - Fix typo	2023-01-04 19:49:14 +08:00
Yifan Xiong	5197cdf5cb	Benchmarks - Support FP8 in BERT models (#446 ) Support FP8 in PyTorch BERT models: * add fp8 hybrid/e4m3/e5m2 in precision arguments * build BERT encoders with `te.TransformerLayer` to repalce `transformers.BertModel` * wrap forward steps with fp8 autocast	2023-01-04 11:12:05 +08:00
Yang Wang	65e433c0c6	Runner: Support `topo-aware` and `k-batch` pattern in 'mpi' mode (#437 ) Description Support the following patterns in `mpi` mode: * `k-batch` * `topo-aware`	2023-01-03 10:28:35 +00:00
Yifan Xiong	616e7a5a5a	Benchmarks - Integrate cublaslt micro-benchmark (#455 ) Integrate cublaslt-gemm micro-benchmark #451.	2023-01-03 08:54:40 +00:00
Yuting Jiang	75573f59da	Benchmarks: Micro benchmarks - Add correctness check in cublas-function benchmark (#452 ) Description Add correctness check in cublas-function benchmark. Major Revision - add python code of correctness check in cublas-function benchmark and test	2023-01-03 14:59:30 +08:00
Yuting Jiang	9dfefce350	Executor - Add stdout logging util module and enable real-time logging flushing in executor (#445 ) Description Add stdout logging util module and enable real-time logging flushing in executor Major Revision - Add stdout logging util module to redirect stdout into file log - enable stdout logging in executor to write benchmark output into both stdout and file `sb-bench.log` - enable real-time log flushing in run_command of microbenchmarks through config `log_flushing` Minor Revision - add log_n_step args to enable regular step time log in model benchmarks - udpate related docs	2022-12-30 09:40:28 +00:00
Yang Wang	7838b6b154	Runner - Support `pair-wise` pattern in `mpi` mode (#447 ) * Extract pair-wise pattern from ib_validation	2022-12-29 08:23:36 +00:00
Yuting Jiang	6583ba2e40	Benchmark: Revision - Add wait time option to resolve mem-bw unstable issue (#438 ) Description Add wait time option to resolve mem-bw unstable issue.	2022-12-14 17:21:02 +08:00
Yang Wang	e4eeda0afd	Runner - support 'pattern' in 'mpi' mode to run tasks in parallel (#430 ) * add mpi-parallels mode * update according to comments * fix and update doc * update * merge into 'mpi' mode * udpate according to comments * fix testcases * fix ansible * regard pattern as field * udpate * fix flake8 version * add flake8 range * remove map-by from host config * udpate comments	2022-11-29 12:30:10 +08:00
Yifan Xiong	1b86503d1e	CLI - Add non-zero return code for `sb [deploy,run]` (#425 ) Add non-zero return code for `sb deploy` and `sb run` command when there're Ansible failures in control plane. Return code is set to count of failure. For failures caused by benchmarks, return code is still set per benchmark in results json file.	2022-11-01 10:46:19 +08:00
Yifan Xiong	d7bb8303fb	CLI - Update version to include revision hash and date (#427 ) Update version to include revision hash and date in "{last tag}+g{git hash}.d{date}" format, here're the examples: * exact tag: 0.6.0 * commit after tag: 0.6.0+gcbb1b34 * commit after tag with local changes: 0.6.0+gcbb1b34.d20221028	2022-10-31 10:44:41 +08:00
Yuting Jiang	3367c4f6cc	Benchmarks - Add support to allow list of custom config string in cudnn-functions and cublas-functions (#414 ) Description Add support to allow list of custom config string in cudnn-functions and cublas-functions.	2022-10-18 09:59:51 +08:00
Yifan Xiong	63e9b2d1bc	Release - SuperBench v0.6.0 (#409 ) Description Cherry-pick bug fixes from v0.6.0 to main. Major Revisions * Enable latency test in ib traffic validation distributed benchmark (#396) * Enhance parameter parsing to allow spaces in value (#397) * Update apt packages in dockerfile (#398) * Upgrade colorlog for NO_COLOR support (#404) * Analyzer - Update error handling to support exit code of sb result diagnosis (#403) * Analyzer - Make baseline file optional in data diagnosis and fix bugs (#399) * Enhance timeout cleanup to avoid possible hanging (#405) * Auto generate ibstat file by pssh (#402) * Analyzer - Format int type and unify empty value to N/A in diagnosis output file (#406) * Docs - Upgrade version and release note (#407) * Docs - Fix issues in document (#408) Co-authored-by: Yang Wang <yangwang1@microsoft.com> Co-authored-by: Yuting Jiang <yutingjiang@microsoft.com>	2022-09-06 18:06:05 +08:00
Yuting Jiang	733860d715	Analyzer - Add support to store values of metrics in data diagnosis (#392 ) Description Add support to store values of metrics in data diagnosis. Take the following rules as example: ``` nccl_store_rule: categories: NCCL_DIS store: True metrics: - nccl-bw:allreduce-run0/allreduce_1073741824_busbw - nccl-bw:allreduce-run1/allreduce_1073741824_busbw - nccl-bw:allreduce-run2/allreduce_1073741824_busbw - nccl-bw:allreduce-run3/allreduce_1073741824_busbw - nccl-bw:allreduce-run4/allreduce_1073741824_busbw nccl_rule: function: multi_rules criteria: 'lambda label:True if min(label["nccl_store_rule"].values())/max(label["nccl_store_rule"].values())<0.95 else False' categories: NCCL_DIS ``` nccl_store_rule will store the values of the metrics in dict and save them into `label["nccl_store_rule"]` , and then rccl_rule can use the values of metrics through `label["nccl_store_rule"].values()` in criteria	2022-08-23 03:25:32 +00:00
Yuting Jiang	10a79c4ea8	Analyzer - Add support for both jsonl and json format in data diagnosis (#388 ) Description Add support for both jsonl and json format in data diagnosis. Major Revision - Add support for both jsonl and json format in data diagnosis Minor Revision - change related doc - add jsonl support in cli	2022-08-22 10:57:00 +08:00
Yuting Jiang	b5c7c85d17	Analyzer: Rename fields in json of data diagnosis to be more readable (#382 ) Description Rename field in data diagnosis to be more readable. Major Revision - rename fields according to diagnosis/metric format Minor Revision - change type of diagnosis/issue_num to be int	2022-08-09 10:03:50 +08:00
Yifan Xiong	9b8df883ae	Gracefully exit when timeout (#383 ) * Gracefully exit when timeout, add corresponding log and return code. * Set minimum timeout to 1 minute and enlarge Ansible timeout.	2022-08-04 13:05:34 +08:00
Yuting Jiang	ec16d42564	Analyzer - Add failure check feature in data diagnosis (#378 ) Description Add failure check feature in data diagnosis. Major Revision - Add failure check rule op to support that if there exists metric_regex not been matched by any metric in result, label as failedtest - Split performance issue and failedtest in categories Minor Revision - replace DataFrame.append() with pd.concat since append() will be removed in later version of pandas	2022-08-01 12:35:35 +08:00
Jie Zhang	ef4d65745b	Support topo-aware IB performance validation (#373 ) * Support topo-aware IB performance validation Add a new pattern `topo-aware`, so the user can run IB performance test based on VM's topology information. This way, the user can validate the IB performance across VM pairs with different distance as a quick test instead of pair-wise test. To run with topo-aware pattern, user needs to specify three required (and two optional) parameters in YAML config file: --pattern topo-aware --ibstat path to ibstat output --ibnetdiscover path to ibnetdiscover output --min_dist minimum distance of VM pairs (optional, default 2) --max_dist maximum distance of VM pairs (optional, default 6) The newly added topo_aware module then parses the topology information, builds a graph, and generates the VM pairs with the specified distance (# hops). The specified IB test will then be running across these generated VM pairs. Signed-off-by: Jie Zhang <jessezhang1010@gmail.com> * Add description about topology aware ib traffic tests Signed-off-by: Jie Zhang <jessezhang1010@gmail.com> * Add unit test to verify generated topology aware config file This commit adds unit test to verify the generated topology aware config file is correct. To do so, four new data files are added in order to invoke gen_topo_aware_config function to generate topology aware config file, then compares it with the expected config file. Signed-off-by: Jie Zhang <jessezhang1010@gmail.com> * Fix lint issue on Azure pipeline Signed-off-by: Jie Zhang <jessezhang1010@gmail.com>	2022-07-26 16:56:19 -07:00
Yang Wang	5d448eedbf	Fix unexpected base conversion when the result value is negative (#377 ) Fix an unexpected result value (`-0.125`) issue in ib traffic benchmark when encountering `-1` in raw output * Check if the value is valid before the base conversion * Add a test case to cover this situation	2022-07-25 15:27:46 +08:00
Yifan Xiong	352ae0c95f	Fix port conflict in ib loopback (#375 ) Fix potential port conflict due to race condition between time-to-check to time-to-use, by binding the port all through. Modify the function to resolve flake8 C901 while keeping the logic same.	2022-07-20 11:30:00 +08:00
Yifan Xiong	b2875179bf	Fix issues in ib validation benchmark (#370 ) Fix several issues in ib validation benchmark: * continue running when timeout in the middle, instead of aborting whole mpi process * make timeout parameter configurable, set default to 120 seconds * avoid mixture of stdio and iostream when print to stdout * set default message size to 8M which will saturate ib in most cases * fix hostfile path issue so that it can be auto found in different cases	2022-07-09 19:57:11 +08:00
Yifan Xiong	e00a8180f6	Support node_num=1 in mpi mode (#372 ) Support `node_num: 1` in mpi mode, so that we can run mpi benchmarks in both 1 node and all nodes in one config by changing `node_num`. Update docs and add test case accordingly.	2022-07-08 09:24:17 +08:00
Yifan Xiong	a94ead34b0	CLI - Support SKU auto detect if running on Azure VM (#365 ) Support SKU auto detect and using corresponding benchmark config if running on Azure VM.	2022-07-05 10:52:39 +08:00
Yifan Xiong	620192a242	Fix issues in ib loopback benchmark (#369 ) Fix several issues in ib loopback benchmark: * use `--report_gbits` and divide by 8 to get GB/s, previous results are MiB/s / 1000 * use the ib_write_bw binary built in third_party instead of system path * update the metrics name so that different hca indices have same metric	2022-06-29 17:53:02 +00:00
Yifan Xiong	bfaa1c837b	Support multiple IB/GPU in ib validation (#363 ) Description Support multiple IB/GPU devices run simultaneously in ib validation benchmark. Major Revisions - Revise ib_validation_performance.cc so that multiple processes per node could be used to launch multiple perftest commands simultaneously. For each node pair in the config, number of processes per node will run in parallel. - Revise ib_validation_performance.py to correct file paths and adjust parameters to specify different NICs/GPUs/NUMA nodes. - Fix env issues in Dockerfile for end-to-end test. - Update ib-traffic configuration examples in config files. - Update unit tests and docs accordingly. Closes #326.	2022-06-24 08:35:20 +00:00
Yifan Xiong	a4937e95c6	Support `sb run` on host directly without Docker (#358 ) Description Support `sb run` on host directly without Docker Major Revisions - Add `--no-docker` argument for `sb run`. - Run on host directly if `--no-docker` if specified. - Update docs and tests correspondingly.	2022-06-14 10:57:01 +08:00
Yuting Jiang	54da021b4d	Analyzer - Fix bugs in data diagnosis (#355 ) Description Fix bugs in data diagnosis. Major Revision - add support to get baseline of the metric which uses custom benchmark naming with ':' like 'nccl-bw:default/allreduce_8_bw:0' - save raw data of all metrics rather than metrics defined in diagnosis_rules.yaml when output_all is True - fix bug of using wrong column index when applying format(red color and percentile) in the excel	2022-06-01 17:12:38 +08:00
Yifan Xiong	6681c72043	Release - SuperBench v0.5.0 (#350 ) Description Cherry-pick bug fixes from v0.5.0 to main. Major Revisions * Bug - Force to fix ort version as '1.10.0' (#343) * Bug - Support no matching rules and unify the output name in result_summary (#345) * Analyzer - Support regex in annotations of benchmark naming for metrics in rules (#344) * Bug - Fix bugs in sync results on root rank for e2e model benchmarks (#342) * Bug - Fix bug of duration feature for model benchmarks in distributed mode (#347) * Docs - Upgrade version and release note (#348) Co-authored-by: Yuting Jiang <v-yutjiang@microsoft.com>	2022-04-29 16:22:55 +08:00
guoshzhao	80dcc8aaec	Benchmarks: Add Benchmark - Add FAMBench based on docker benchmark (#338 ) Description Integrate FAMBench into superbench based on docker implementation: https://github.com/facebookresearch/FAMBench The script to run all benchmarks is: https://github.com/facebookresearch/FAMBench/blob/main/benchmarks/run_all.sh	2022-04-11 15:31:07 +08:00
Yuting Jiang	8dc19ca4af	CLI - Integrate output all nodes diagnosis results (#339 ) Description Integrate output all nodes diagnosis results.	2022-04-11 13:42:04 +08:00
Yuting Jiang	55b0f9d239	Analyzer: Add Feature - Output results of all nodes in data diagnosis (#336 ) Description Output results of all nodes in data diagnosis.	2022-04-10 18:57:15 +08:00
Yuting Jiang	f15da60b2b	CLI - Integrage result summary and update output format of data diagnosis (#335 ) Description Integrage result summary and update output format of data diagnosis. Major Revision - integrage result summary - add md and html format for data diagnosis	2022-04-08 18:48:43 +08:00
guoshzhao	6d895da83c	Benchmarks: Add Feature - Provide option to save raw data into file. (#333 ) Description Use config `log_raw_data` to control whether log the raw data into file or not. The default value is `no`. We can set it as `yes` for some particular benchmarks to save the raw data into file, such as NCCL/RCCL test.	2022-04-01 16:26:09 +08:00
Yuting Jiang	84fed1ce18	Analyzer: Add feature - Add result summary in excel,md,html format (#320 ) Description Add result summary in excel,md,html format. Major Revision - Add ResultSummary class to support result summary in excel,md,html format. - Abstract RuleBase class for common-used functions in DataDiagnosis and ResultSummary.	2022-03-24 15:32:01 +08:00
rafsalas19	ff51a3cee9	Benchmarks: Add Feature - Add GPU-Burn as microbenchmark (#324 ) Description Modifications adding GPU-Burn to SuperBench. - added third party submodule - modified Makefile to make gpu-burn binary - added/modified microbenchmarks to add gpu-burn python scripts - modified default and azure_ndv4 configs to add gpu-burn	2022-03-16 16:20:11 +08:00
Yuting Jiang	b3c95f1827	Analyzer - Add md and html output format for DataDiagnosis (#325 ) Description Add md and html output format for DataDiagnosis. Major Revision - add md and html support in file_handler - add interface in DataDiagnosis for md and HTML output Minor Revision - move excel and json output interface into DataDiagnosis	2022-03-15 18:04:11 +08:00
Yuting Jiang	1ec055e1c2	Analyzer: Revise - Abstract RuleBase from DataDiagnosis (#321 ) Description Abstract RuleBase from DataDiagnosis.	2022-03-07 17:25:07 +08:00
Yuting Jiang	97ed12f97f	Analyzer: Add Feature - Add multi-rules feature for data diagnosis (#289 ) Description Add multi-rules feature for data diagnosis to support multiple rules' combined check. Major Revision - revise rule design to support multiple rules combination check - update related codes and tests	2022-02-20 16:59:38 +08:00
Ziyue Yang	6cdf759543	Benchmarks: Revise Code - Eliminate NUMA binding for device-to-device tests in gpu_copy (#302 ) Description This commit remove NUMA binding for device-to-device tests because NUMA doesn't affect performance, and revise benchmark metrics accordingly.	2022-02-09 20:30:42 +08:00
Ziyue Yang	682b2c120d	Benchmarks: Revise Code - Make data checking in gpu_copy optional (#301 ) This commit makes data checking in gpu_copy optional, because it will take too long time if message size is large.	2022-02-08 10:59:27 +08:00

1 2 3 4

162 Коммитов