Other
AI Platform Engineer
Toronto , OntarioHybridContract12 monthsPosted August 26, 2026
Job Description
We are looking for an experienced AI Platform Engineer to help drive enterprise AI platform enablement at the intersection of GenAI engineering, cloud technology, governance, and user enablement.
You will help build and enable secure, scalable AI capabilities while partnering with Information Security, Risk, Legal, Compliance, engineering, and business teams to translate enterprise controls into practical technical solutions. You will also support AI platform users and help manage AI consumption, including tokens, credits, quotas, and platform costs.
This is a hands-on role for someone who combines strong software engineering fundamentals with practical experience in GenAI platforms, LLM applications, AWS, Agentic AI, MCP, and RAG architectures.
What You Will Be Doing
Lead AI platform enablement and onboard users, teams, and use cases across platforms including OpenAI, Gemini, Anthropic, Amazon Bedrock, LLM gateways, and agentic frameworks
Partner with InfoSec, Risk, Compliance, Legal, and engineering teams to define and implement platform controls and guardrails
Implement technical controls including IAM policies, content filters, logging/monitoring, rate limits, and data-access controls
Manage token budgets, credits, quotas, rate limits, and AI platform costs, including dashboards, anomaly detection, and automated alerts
Support users across AI platform capabilities, including agents, skills, MCP integrations, prompt-based applications, APIs, and responsible AI usage
Build scalable LLM pipelines, agentic workflows, MCP integrations, and Graph/RAG architectures using Python and AWS-native tooling
Deliver cloud-native solutions using EKS, Lambda, Fargate, Glue, and Athena
Automate provisioning, controls, policy checks, and cost guardrails using Infrastructure as Code
Maintain platform documentation, runbooks, onboarding guides, and audit/compliance evidence
Own end-to-end delivery of platform initiatives using Agile Scrum/Kanban practices and contribute as an SME to the AI platform roadmap
What You Bring
6+ years of engineering experience, including 1–2+ years in AI platform, cloud platform, or emerging-tech enablement
Hands-on experience with GenAI models (GPT, Claude, Gemini, LLaMA), prompt engineering, Agentic AI, MCP, Graph/RAG, and LLM gateway/proxy patterns
Strong Python skills, including NumPy, Pandas, and Boto3
Strong AWS experience across services including AgentCore, Bedrock, EC2, ELB/GLB/NLB, EKS, Fargate, Lambda, Athena, Glue, and Lake Formation
Experience implementing enterprise security and governance controls in partnership with InfoSec, Risk, and Compliance
Experience with Terraform, Puppet, Docker, Infrastructure as Code, and containerized deployments
Experience with Vector/Graph databases such as Weaviate, Milvus, PGVector, Neo4j, and Neptune
Experience with automated testing/evaluation tools including Ragas, Playwright, Selenium, and Zephyr
Strong knowledge of DevSecOps, SDLC, Agile Scrum/Kanban, JIRA, Confluence, and JIRA Align
Strong stakeholder management, communication, and user-support skills
Nice to have: Experience with QuickSight or Tableau for usage and cost reporting, along with knowledge of financial markets and enterprise data systems.
Requirements
- Lead AI platform enablement and onboard users, teams, and use cases across platforms including OpenAI, Gemini, Anthropic, Amazon Bedrock, LLM gateways, and agentic frameworks
- Partner with InfoSec, Risk, Compliance, Legal, and engineering teams to define and implement platform controls and guardrails
- Implement technical controls including IAM policies, content filters, logging/monitoring, rate limits, and data-access controls
- Manage token budgets, credits, quotas, rate limits, and AI platform costs, including dashboards, anomaly detection, and automated alerts
- Support users across AI platform capabilities, including agents, skills, MCP integrations, prompt-based applications, APIs, and responsible AI usage
- Build scalable LLM pipelines, agentic workflows, MCP integrations, and Graph/RAG architectures using Python and AWS-native tooling
- Deliver cloud-native solutions using EKS, Lambda, Fargate, Glue, and Athena
- Automate provisioning, controls, policy checks, and cost guardrails using Infrastructure as Code
- Maintain platform documentation, runbooks, onboarding guides, and audit/compliance evidence
- Own end-to-end delivery of platform initiatives using Agile Scrum/Kanban practices and contribute as an SME to the AI platform roadmap
- 6+ years of engineering experience, including 1–2+ years in AI platform, cloud platform, or emerging-tech enablement
- Hands-on experience with GenAI models (GPT, Claude, Gemini, LLaMA), prompt engineering, Agentic AI, MCP, Graph/RAG, and LLM gateway/proxy patterns
- Strong Python skills, including NumPy, Pandas, and Boto3
- Strong AWS experience across services including AgentCore, Bedrock, EC2, ELB/GLB/NLB, EKS, Fargate, Lambda, Athena, Glue, and Lake Formation
- Experience implementing enterprise security and governance controls in partnership with InfoSec, Risk, and Compliance
- Experience with Terraform, Puppet, Docker, Infrastructure as Code, and containerized deployments
- Experience with Vector/Graph databases such as Weaviate, Milvus, PGVector, Neo4j, and Neptune
- Experience with automated testing/evaluation tools including Ragas, Playwright, Selenium, and Zephyr
- Strong knowledge of DevSecOps, SDLC, Agile Scrum/Kanban, JIRA, Confluence, and JIRA Align
- Strong stakeholder management, communication, and user-support skills
Interested in this position?
Apply now and our recruitment team will be in touch with you shortly.
Apply for This Position