Add ChakraTS (YHat Labs) API wrapper and fev-bench results - #192
Open
shubham-forum wants to merge 1 commit into
Open
shubham-forum wants to merge 1 commit into
shubham-forum wants to merge 1 commit into
Conversation
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #, if available: None
Description of changes:
Add the
chakrats_apiwrapper and pinned dependencies for ChakraTS, a hosted forecasting model from YHat Labs (https://huggingface.co/yhatlabs/ChakraTS). No weights are distributed; the wrapper reads the API key fromYHAT_API_KEY, no credentials are included. This submission is for a closed-source API model as discussed with Oleksandr.Add results for all 100 fev-bench tasks in
benchmarks/fev_bench/results/chakrats_api.csv: SQL skill score 48.1 (seasonal-naive normalised). The results are forchakra-ts-fev, the fev-bench evaluation configuration of ChakraTS: no part of it was pretrained on any fev-bench dataset, sotrained_on_datasetsis empty and no task is flagged.Add the
models.yamlentry (closed-api, zero-shot, commercial use permitted).The wrapper sends each window's items with all target columns, past covariates and known covariates, 100 items per call, and never mixes windows in one call.
Validation:
test/test_fev_bench_submissions.pypasses for the new file and metadata.The result CSV has 100 unique tasks and complete SQL, MASE, WAPE and WQL values.
The wrapper run against the API reproduces the submitted rows within 0.1% on every task we checked. We will send the maintainers an evaluation key privately.
-t 3runs the three ProEnFo tasks in about 5 minutes,-t 8about 45 minutes, the full benchmark about few hours. The server scales to zero when idle, so the first call can take a couple of minutes; the wrapper retries.We can also share the complete outputs of our own run privately, including the final forecasts for all 100 tasks, if that helps the review.
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.