Text this: Bridging semantics and vision: text-guided feature alignment for few-shot object detection