waypointjobs

Cisco Systems, Inc.

Software Engineer

san jose, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Cisco Systems, Inc. lists this Software Engineer opportunity in san jose, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

The application window is expected to close on: 10/29/2026We run the platform that serves foundational models to Cisco IT. Our Foundational Model Service, gives engineering teams across the company access to small, large, and embedding models. youWe serve those models on Kubernetes clusters, with Nim, Vllm and other runtimes. Beyond serving, We benchmark, evaluate, monitor, and release new models as improvements and demand warrant. We use published model artifacts where possible, refining or rebuilding them for compatibility, performance, or quality. Our customers depend on the platform under a 99.9% uptime SLA, and we build and operate accordingly.

As a Senior Software Engineer, you'll guide model serving, runtime tuning, accelerator optimization, and release evaluation. You'll lead incident response and improve reliability, performance, and operability. You'll also use agentic workflows to identify problems earlier and automate remediation.

You'll follow Cisco Design Thinking Principles, simplify user experience, apply secure coding practices, and protect privacy. You'll work across design, product, and engineering to improve customer solutions, documentation, development practices, and production reliability.

This role calls for production experience in AI infrastructure, model serving, evaluation, and related services. We're looking for a self starter who works independently, turns ambiguous problems into plans, mentors engineers, and raises technical standards. The platform evolves with new runtimes, accelerators, and models. On call is shared across the team.

Responsibilities

Contribute to model serving direction and roadmap, including runtime selection and tuning across vLLM, NVIDIA NIM, and other runtimes.

Guide quantization and accelerator optimization across GPU vendors, validating performance, quality, capacity, and cost with data

Develop and enhance platform services, APIs, gateways, and operational tooling around model inference

Evolve evaluation, benchmarking, and load test frameworks that gate model releases against service level objectives

Define model promotion criteria across quality, safety, latency, throughput, and resource use

Evaluate how fine tuning, distillation, and quantization affect production behavior

Shape routing and capacity behavior, including prefix caching, KV aware routing, and prefill and decode separation

Improve model registry, packaging, evaluation, release, and development workflows using Infrastructure as Code, GitHub Actions, and agentic workflows

Refine model artifacts when runtime compatibility or performance requires changes

Develop observability that shows platform health, model performance, capacity, customer adoption, and usage

Monitor production, serve as an escalation point for on call issues, lead postmortems and root cause analyses, and drive durable improvements

Coordinate across customer, product, design, and engineering teams to gather input, forecast capacity, track milestones, and guide platform direction

Apply AI to platform operations through anomaly detection, automated remediation, and predictive operations

Lead features and projects from technical design through completion, working with minimal guidance and driving results through delegation and review

Write clean code and unit tests independently, and review code for quality, threat models, scale, reliability, and release velocity

Act as a technical resource, mentor engineers, run design reviews, and share knowledge across teams

Create technical designs, runbooks, user documentation, project updates, and remediation plans

Requirements

7 or more years of related systems, platform, or software engineering experience, or equivalent practical experience, with solid knowledge across related technologies

A production background developing and operating AI infrastructure

A solid understanding of LLM, SLM, embedding, and reranker model internals, including context length, batching, token throughput, and memory use

Hands on work serving models on inference runtimes including vLLM, NVIDIA NIM, or Triton

Proven results building services around models, including the APIs, gateways, and operational tooling that make them consumable

Depth in evaluating and benchmarking models, and using the results to make release decisions

Familiarity moving model artifacts through evaluation, optimization, packaging, and production serving

Deep understanding of supervised fine tuning, parameter efficient fine tuning, distillation, and quantization

Command of distributed GPU training concepts, mixed precision, and parallelism strategies

Production Kubernetes work running GPU workloads at scale

Linux administration and troubleshooting

Programming in Python or Go

Fluency with CI/CD pipelines and Infrastructure as Code, for example Terraform or Ansible

mastery of monitoring and observability tooling, including Prometheus, Grafana, or Splunk

A track record of mentoring engineers and setting technical standards

A demonstrated pattern of learning new systems and the initiative to lead unfamiliar work

Ability to take part in an on call rotation for a service the company depends on

Clear written communication and the habit of documenting what you develop

Preferred Knowledge and Experience

Work with AMD accelerators and ROCm alongside NVIDIA Cuda

Exposure to distributed inference, disaggregated serving, or KV cache aware routing

Familiarity with evaluation frameworks including lm-eval or DeepEval, and with safety and capability suites

A background evaluating RAG or agent systems

Time spent with API gateways, ingress, or load balancing, for example Envoy, APISIX, or NGINX

Fluency with GitOps and Helm, particularly ArgoCD

Capacity planning, traffic pattern understanding, and cost optimization for GPU fleets

Disaster recovery for stateful platform services

Contributions to open source inference or evaluation projects

Certified Kubernetes Administrator (CKA) or an equivalent cloud certification

Training or fine tuning transformer models with PyTorch and Hugging Face

Practical use of LoRA, QLoRA, and PEFT

Distributed training with PyTorch FSDP, DeepSpeed, or comparable frameworks

Model registries and experiment tracking systems, for example MLflow or Weights and Biases

Multi node GPU training and collective communication libraries

Education

Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent practical experience

Why Cisco?

At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations in the AI era - and beyond. We've been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you'll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.

We are Cisco, and our power starts with you.

Message to applicants applying to work in the U.S. and/or Canada:

The starting salary range posted for this position is $167,700.00 to $245,200.00 and reflects the projected salary range for new hires in this position in U.S. and/or Canada locations, not including incentive compensation*, equity, or benefits. Individual pay is determined by the candidate's hiring location, market conditions, job-related skillset, experience, qualifications, education, certifications, and/or training. The full salary range for certain locations is listed below. For locations not listed below, the recruiter can share more details about compensation for the role in your location during the hiring process.

U.S. employees are offered benefits, subject to Cisco's plan eligibility rules, which include medical, dental and vision insurance, a 401(k) plan with a Cisco matching contribution, paid parental leave, short and long-term disability coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible to receive grants of Cisco restricted stock units, which vest following continued employment with Cisco for defined periods of time.

U.S. employees are eligible for paid time away as described below, subject to Cisco's policies:

10 paid holidays per full calendar year, plus 1 floating holiday for non-exempt employees

1 paid day off for employee's birthday, paid year-end holiday shutdown, and 4 paid days off for personal wellness determined by Cisco

Non-exempt employees** receive 16 days of paid vacation time per full calendar year, accrued at rate of 4.92 hours per pay period for full-time employees

Exempt employees participate in Cisco's flexible vacation time off program, which has no defined limit on how much vacation time eligible employees may use (subject to availability and some business limitations)

80 hours of sick time off provided on hire date and each January 1st thereafter, and up to 80 hours ofunused sick timecarried forwardfrom one calendar yearto the next

Additional paid time away may be requested to deal with critical or emergency issues for family members

Optional 10 paid days per full calendar year to volunteer

For non-sales roles, employees are also eligible to earn annual bonuses subject to Cisco's policies.

Employees on sales plans earn performance-based incentive pay on top of their base salary, which is split between quota and non-quota components, subject to the applicable Cisco plan. For quota-based incentive pay, Cisco typically pays as follows:

.75% of incentive target for each 1% of revenue attainment up to 50% of quota;

1.5% of incentive target for each 1% of attainment between 50% and 75%;

1% of incentive target for each 1% of attainment between 75% and 100%; and

Once performance exceeds 100% attainment, incentive rates are at or above 1% for each 1% of attainment with no cap on incentive compensation.

For non-quota-based sales performance elements such as strategic sales objectives, Cisco may pay 0% up to 125% of target. Cisco sales plans do not have a minimum threshold of performance for sales incentive compensation to be paid.

The applicable full salary ranges for this position, by specific state, are listed below:

New York City Metro Area:

$167,700.00 - $282,000.00 Non-Metro New York state & Washington state:

$149,100.00 - $250,900.00 * For quota-based sales roles on Cisco's sales plan, the ranges provided in this posting include base pay and sales target incentive compensation combined.

** Employees in Illinois, whether exempt or non-exempt, will participate in a unique time off program to meet local requirements.

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Embedded Systems Developer

laurel, MD

See pay details in description

Description Do you love working on a motivated team to solve complex problems in innovative ways? Do you enjoy creating embedded prototypes i…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Oracle E-Business Suite (EBS) Developer

laurel, MD

See pay details in description

Description Are you a skilled problem-solver who blends technical expertise with creativity to build effective and efficient solutions? If s…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Software Engineer

laurel, MD

See pay details in description

Description Are you passionate about building solutions for our greatest national security challenges? Are you searching for engaging work wi…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Reverse Engineer and AI Workflow Developer

laurel, MD

See pay details in description

Description Are you passionate about reverse engineering complex software, firmware, and hardware? Do you enjoy developing agentic AI workflo…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

AFSIM Developer

laurel, MD

See pay details in description

Description Are you passionate about building high-fidelity simulations that shape the future of U.S. Space Force capabilities and force design …

Listing review due 2026-10-06View job