Predicting future frames and improving inter-frame prediction are ongoing challenges in the field of video streaming. By creating a novel framework called STreamNet (Spatial-Temporal Video Coding), fusing bidirectional long short-term memory with temporal convolutional networks, this work aims to address the issue at hand. The development of STreamNet, which combines spatial hierarchies with local and global temporal dependencies in a seamless manner, along with sophisticated preprocessing, attention mechanisms, residual learning, and effective compression techniques, is the main contribution. Significantly, STreamNet claims to provide improved video coding quality and efficiency, making it suitable for next-generation networks. STreamNet has the potential to provide reliable and optimal streaming in high-demand network environments, as shown by preliminary tests that show a performance advantage over existing methods.