Priyanshu Yadav

● open to roles · IST

← work

Employer work · private code

An HR assistant that answers from policy and acts in the system

Employees ask HR questions in plain language. The assistant answers from the company’s own policies with citations, and when a question needs live data it calls the HRMS: leave balances, attendance, raising a ticket. It can only do what the person asking is allowed to do.

Role
AI Architect at RazorInfoTech
Year
2026
Built with
Python, LangChain, LangGraph, RAG, Tool calling, vLLM, Langfuse
2 pathsgrounded policy answers, and live actions through tools
3 toolsleave balances, attendance, ticket creation
Per roleaccess checks before any tool runs
An employee's question goes to an LLM router. One path retrieves policy passages and cites them. The other path proposes a tool call, which passes a permission check before an HRMS tool runs. Both paths end at a schema check before the reply. A denied tool call returns an explanation and makes no call. every hop traced · Langfuse · tokens · cost · latencyEmployeeHRMS chatRouterLLM · served by vLLManswer | tool_callRetrievepolicy index · top-k · citePermission checkrole → allowed toolsHRMS toolsleave · attendance · ticketdenied → explain, no callSchema checkvalidate → reply

The problem

HR teams answer the same policy questions all day, and employees open the HRMS mostly to look up one number. A chat box that only talks is not much help. It has to read the real policy and touch the real system, and it must never show one employee another employee’s data.

What I built

  • A retrieval pipeline over HR policy documents, so answers are grounded and carry citations back to the source passage.
  • Function-calling interfaces for live HRMS actions: leave balances, attendance and ticket creation.
  • Role-based access control in front of every tool. The model can request an action, and the system decides whether this user may run it.
  • Structured output validation, so a malformed model response fails safely before it reaches the HRMS.

Running it

Open-source models are served on cloud GPUs with vLLM for low-latency, high-volume traffic. Langfuse traces every request, which gives token and cost tracking, latency numbers and an evaluation loop for answer quality. Releases go through CI/CD with version control and monitoring.