In recent years, with increase in the use of internet the multimedia contents on it have rapidly increased. Users may need to go through a video in a top down manner i.e. browsing the videos, or in bottom up manner i.e. retrieving specific information from videos. They may also want to go through the summary or through the highlights of the videos. This has necessitated the need to handle multimedia resources effectively. This paper proposes an automatic method for aligning scripts of lecture videos with captions. Alignment is needed to extract time information from captions and insert it in the scripts, to create index of the videos. No alignment work has been previously done in lecture videos domain. Alignment methods proposed for other type of videos are not applicable for lecture videos because, different similarity techniques behave differently on different types of datasets. The proposed method uses transcripts of lecture videos, SRT file of captions available along with lecture videos and captions generated from auto-caption generation feature of YouTube. The captions and scripts are then aligned using a dynamic programming technique. No such work has been previously done for lecture videos. Most important aspect of alignment is similarity measure. In the proposed work we have used three similarity measures cosine, jaccard, and dice. A comparative analysis of these measures is given in the paper. We also use a large lexical database of English words known as WordNet for word-to-word similarity. The experimental result shows comparison of various similarity techniques and YouTube captions.