Replace the SageMaker Python SDK with direct boto3 calls - #281
Conversation
|
Job PR-281-c4683e8 is done. |
…ures to their jobs - Replace role/vpc_config/kms_key/tags constructor args with backend=SageMakerConfig(...), shared by cloud predictors and foundation models; add explicit region and split output_kms_key/volume_kms_key. - Rename sagemaker_overrides to backend_overrides and accept it on foundation-model predict. - Bind JobPredictionFuture to the submitted job and upload fit inputs under a per-job prefix so later submissions don't overwrite earlier ones.
…ker-core resources sagemaker-core resource classes route every call through a process-wide client that ignores the session argument, which breaks per-object region/credentials. Send requests through the session's own sagemaker / sagemaker-runtime clients instead. - Build requests in SageMaker API / boto3 PascalCase; backend_overrides fields use the same format as the AWS API reference. Override keys stay the boto3 method names. - SageMakerConfig.vpc_config and inference_config keep their snake_case keys and are mapped explicitly; unknown keys raise before any resource is created. - Wait for endpoints with botocore's endpoint_in_service waiter. - Remove the sagemaker-core session-binding and acronym-serialization workarounds. - Unit tests validate generated requests against botocore's service model.
|
Job PR-281-7e43a3d is done. |
The remaining uses were thin helpers around boto3. Replace them with small local implementations so all AWS calls go through the backend's own boto3 session. - AwsSession wraps a boto3.Session with cached sagemaker / sagemaker-runtime / s3 clients plus upload_data / download_data. - get_execution_role derives the role from the caller's STS identity and resolves its path via IAM (falling back to service-role/ for SageMaker console roles); IAM users get an actionable error instead of a guess. - repack_model_with_serving_code repacks the model tarball directly. - Move sagemaker_timestamp / unique_name_from_base to utils.misc and drop the serializer/deserializer base classes.
|
Job PR-281-ca921eb is done. |
Keep the public API as close to master as possible so the PR only swaps the SageMaker SDK for direct boto3 calls. - Restore the `role=` constructor arg and `custom_image_uri`; remove SageMakerConfig, vpc/kms/user tags and the environment / use_spot_instances / max_wait args (all reachable via backend_overrides). - Revert unrelated changes: prediction-future binding, per-job upload prefix, docs/API surface tweaks, internal backend refactors.
|
Job PR-281-7b0778f is done. |
- Replace predict()/predict_proba()'s download / persist / save_path with predictions_path, the S3 prefix the batch transform job writes to. With wait=True results are always loaded and returned, as in fit_predict(). - Inline resolve_image_uri: use custom_image_uri or retrieve_image_uri().
|
Job PR-281-1d51cdd is done. |
- Clean up the model and endpoint config when deploy() fails partway, and give each deploy a unique endpoint config name so a leftover config can't block a redeploy. - Batch transform: delete only the model this call created, and register the job under its actual (possibly overridden) name. - Pass training code as a `code` input channel, as SDK v2 did, so training works with EnableNetworkIsolation. - attach_job() raises again when the training job did not complete.
- backend_overrides can't set fields that link created resources (ProductionVariants / their ModelName, EndpointConfigName on the endpoint, ModelName on the transform job), so cleanup only deletes our own resources. - Validate inference_mode before creating anything. - Rollback deletes log failures instead of masking the original error.
|
Job PR-281-d56f415 is done. |
|
Job PR-281-025258c is done. |
Draft. Removes the SageMaker Python SDK dependency (
sagemaker>=2.240,<3). AutoGluon-Cloud now builds SageMaker API requests itself and sends them through boto3.Dependency
sagemakeris dropped; there is no replacement (nosagemaker-core). Onlyboto3is needed.Internals
CreateTrainingJob/CreateModel/CreateEndpointConfig/CreateEndpoint/CreateTransformJobrequests. The v2 Estimator/Model/Predictor subclasses are gone.AwsSession(inutils/aws_utils.py) holds the boto3 session and its sagemaker / sagemaker-runtime / s3 clients, plus S3 upload/download helpers. All calls go through the backend's own session.sourcedir.tar.gzand passed to the container as acodeinput channel (soEnableNetworkIsolationworks); serving code is repacked intomodel.tar.gzundercode/.SagemakerEndpointand theEndpointbase class are removed.API changes (breaking)
backend_kwargs(which held SDK-specific dicts) is replaced bybackend_overrides: raw SageMaker request fields keyed by boto3 method name (create_training_job,create_model,production_variant,create_endpoint_config,create_endpoint,create_transform_job). They are deep-merged over the generated requests. Fields that link the created resources (e.g.ProductionVariants,ModelNameon the variant / transform job,EndpointConfigNameon the endpoint) can't be overridden, so cleanup only ever deletes resources AutoGluon-Cloud created.predict()/predict_proba()takepredictions_path(S3 prefix for the batch transform output) likefit_predict(), and always return the loaded results whenwait=True. Thedownload/persist/save_pathoptions (previously keys ofbackend_kwargs) are removed.backend_kwargs,model_kwargs,deploy_kwargs,transformer_kwargs, …) raise an error naming their replacement.attach_endpoint()takes an endpoint name;detach_endpoint()returns the endpoint name.entry_point/source_dirin the old SDK kwargs) are no longer supported and raise an explicit error; usecustom_image_urito customize the container.Everything else is unchanged from master (
role=,custom_image_uri, region resolution, tags).Migration
fit(..., backend_kwargs={"autogluon_sagemaker_estimator_kwargs": {...}})fit(..., backend_overrides={"create_training_job": {...}})with CreateTrainingJob fieldsdeploy(..., backend_kwargs={"model_kwargs": {...}, "deploy_kwargs": {...}})deploy(..., backend_overrides={"create_model": {...}, "production_variant": {...}, "create_endpoint_config": {...}})predict(..., backend_kwargs={"transformer_kwargs": {...}})predict(..., backend_overrides={"create_transform_job": {...}})predict(df, backend_kwargs={"download": True, "save_path": p})predict(df).to_csv(p); results are always returned whenwait=Truepredict(df, backend_kwargs={"download": False})predict(df, wait=False), thenget_batch_inference_job_info()for the S3 result pathpredict(df, backend_kwargs={"transform_kwargs": {"output_path": uri}})predict(df, predictions_path=uri)detach_endpoint()returned anEndpointobjectattach_endpoint(name)instance_type="local"Examples:
Known gaps / follow-ups
Tagsviabackend_overridesreplaces the defaultautogluon-cloud-*tags, since lists are not merged.backend_overrides.CloudPredictor.deploy()returning an endpoint object, matchingFoundationModel.deploy().