Taro Logo

Deep Learning Performance Architect Interview Experience - Santa Clara, California

October 1, 2024
Negative ExperienceNo Offer

Process

The process took around two months, during which I had two interviews via Teams. These interviews, spread over two months, included one with a project member and one with the project lead. Basic questions about my resume and fundamental architecture were asked.

After that, I proceeded to the panel round, where I interviewed with four team members.

The first panel interview went well. Although I struggled when performing a technical question by hand, I still arrived at the correct answer.

The second interviewer did not ask many technical questions. Instead, they focused on my resume, asking about topics I hadn't worked on. It felt odd; after stating I had experience with x, y, and z, the interviewer asked about something tangential, which, based on my impression from other interviews, they didn't seem to care about. I provided a correct answer but with flawed reasoning. However, I answered the other technical questions related to my experience adequately.

The third interview was a complete mess. The interviewer did not clearly define any of the problems. Imagine being repeatedly asked to perform the operation 'what's 5 and 6?' and, after asking for clarification, receiving 'what's 5 and 6 together?' This forced me to guess each time whether the operation was addition (11) or multiplication (30).

I should have recognized the issue when I referenced a common operation in PyTorch/NumPy/Tensorflow, and the interviewer seemed clueless about the notation.

The final interview was good but fundamentally suffered from the same problems as the others. They asked if I knew 'tensor parallelism.' My immediate thought was, 'What a strange name; you fundamentally have three tensors—input, output, and weight—and only one of them is parallel.'

While I felt this was my strongest interview, as I had become accustomed to asking for definitions and taking calculated risks if they refused to provide them, it still felt like pulling teeth when discussing topics I had been working on closely for the past two years.

I have extensive experience with these problems in the accelerator space and assumed that even though this job was GPU-based, I would be well-prepared.

The accelerator space seems to define these terms more clearly, discussing them in terms of concepts like 'fine-grained parallelism' or 'coarse-grained parallelism' (referring to proximity to compute, as seen in Tangram from Stanford). They also use terms like 'weight stationary,' 'output stationary,' 'input stationary,' or 'row stationary' (as explored in work by Joel Emer and Vivienne Sze) to define which part of the tensor is moving (though this is primary; secondary movements don't have a defined order but can be notated for complete movement). Additional work even incorporates high-level languages or representations to define these movements (see Maestro from GATech).

Transitioning from that space to this feels like speaking a completely different language, where every word sounds the same but has no consistent meaning. To be clear, I am willing to learn these translations, words, or definitions. However, being asked to do so on the spot without preparation strikes me as unreasonable.

It was frustrating to constantly fight to translate loosely defined terms or poorly conceived problems throughout the panel interviews. To answer, I had to be entirely reactive.

I don't believe any role would actually require this. Typically, one would be expected to sit down and understand concepts before simulating them. However, this process involved a mix of poorly formulated problems and definitions that often only made sense to the interviewers.

I don't believe they effectively tested my ability to do the job. At best, they tested if I could perform at their level from day one, which seems unreasonable.

I would find it difficult for anyone to pass the same process I underwent without additional information about the terminology, the types of problems being asked, clearer problem definitions, or some form of nepotism.

Questions

How to analyze deep learning primitives.

Was this helpful?

Interview Statistics

The following metrics were computed from 1 interview experience for the Nvidia Deep Learning Performance Architect role in Santa Clara, California.

Success Rate

0%
Pass Rate

Nvidia's interview process for their Deep Learning Performance Architect roles in Santa Clara, California is extremely selective, failing the vast majority of engineers.

Experience Rating

Positive0%
Neutral0%
Negative100%

Candidates reported having very negative feelings for Nvidia's Deep Learning Performance Architect interview process in Santa Clara, California.

Nvidia Work Experiences