Positron Runs on SageMaker AI
- •Positron now runs on Amazon SageMaker AI for governed data science workflows in Studio Spaces
- •AWS walkthrough used a synthetic 50,000-loan portfolio with Athena, R, Python, Shiny, and Quarto
- •Recorded run produced AUC 0.834, 48,500 scored loans, and an InService real-time endpoint
Amazon Web Services said on September 21, 2026, that Positron, Posit’s integrated development environment for data science, now runs on Amazon SageMaker AI so data science teams can use governed data access, R analysis, Python model development, deployment, application development, and reporting from one browser-based SageMaker Studio Space. Posit publishes a container image definition for Positron built on the Amazon SageMaker Distribution image; administrators build it, push it to Amazon Elastic Container Registry, register it with SageMaker AI, and attach it to a Studio domain. Data scientists then select Positron when creating a Space and open the IDE directly in Studio.
AWS described a recorded walkthrough using a synthetic 50,000-loan portfolio, with Amazon S3 storing source data, the AWS Glue Data Catalog registering it, Amazon Athena querying it, R validating features, and Python training an XGBoost classifier. The setup required a Posit license grant, access to the Posit-published Positron image definition, administrator permissions for Amazon ECR and SageMaker Studio custom images, a Space execution role with Athena and AWS Glue Data Catalog access, S3 and Athena result locations, Amazon Bedrock model access in the same AWS Region when using Posit Assistant, and an ml.t3.xlarge instance or larger.
The workflow began in a SageMaker Studio Space where Positron provided the project explorer, editor, R and Python sessions, Variables pane, plots, terminal, and application preview under the Space execution role. Posit Assistant identified credit_risk_blog.loan_tape_source in the AWS Glue Data Catalog and prepared a read-only Athena query that returned five sample rows across six fields, scanned 2.18 MiB, and finished in under one second. When Posit Assistant used Amazon Bedrock as its model provider, AWS said no separate model-provider API key was required if authentication resolved through the environment’s AWS credentials; the captured session showed 6,657,942 tokens, including 6,118,411 cache-read and 462,905 cache-write tokens, an estimated cost of $6.319, and 92.5 percent cache efficiency.
The data science run profiled the 50,000-loan table in Athena, finding 1,500 records with missing income, 1,015 defaults, and an overall default rate of 2.03 percent. R loaded the 50,000 rows, created debt-to-income and log-income features, and left 48,500 loans for modeling and scoring after excluding the 1,500 incomplete records. Python then trained an XGBoost classifier on 40,000 rows and three model features; the held-out evaluation produced an AUC of 0.834 and a 12.3 percent observed default rate in the highest-risk decile.
The recorded workflow wrote predicted probabilities and risk deciles for 48,500 loans to Parquet, registered the results as credit_risk_blog.scored_loans in Athena, and created a SageMaker AI model, endpoint configuration, and real-time endpoint that reached InService. A Shiny for Python application invoked the endpoint under the Space execution role, used the training feature definitions on synthetic applicant data, and displayed the returned probability of default. The run ended with a Quarto report connecting the Athena source, data-quality findings, R validation, Python model, scored output, SageMaker AI endpoint, and Shiny application.
AWS said the architecture separates an administrator path from a data science path: administrators manage the custom image, Amazon ECR repository, SageMaker AI image and version, licensing, permissions, and domain attachment, while data scientists use R, Python, Quarto, Posit Database Drivers, and optional Posit Assistant inside the Space. AWS also stated the dataset and applicant payloads were synthetic and that the run did not establish model fairness, calibration, lending suitability, production latency, load behavior, monitoring, or regulatory compliance. The AUC and decile results came from one held-out split, lower deciles were not strictly monotonic, and production use still depends on customer controls for security, governance, validation, networking, logging, patching, and operations.